The Agentic Security Triad: A New Hypothesis for Assuring Advanced Artificial Intelligence Remains Safe for Humanity
Can technology really become dangerous enough to destroy humanity?
We have already crossed that bridge once.
When scientists learned how to release the energy contained within the atomic nucleus, humanity acquired a technological capability powerful enough to threaten civilization itself. Nuclear weapons did not end technological progress, but they permanently changed the relationship between innovation and consequence. Once humanity possessed them, we could no longer pretend that every technological failure would be local, recoverable, or even survivable.
Artificial intelligence raises a different but increasingly serious version of the same question. We do not know whether artificial superintelligence will be created, precisely what form it would take, or whether it would pose an existential danger. Those remain subjects of substantial scientific disagreement. But artificial intelligence is already moving beyond systems that merely answer questions. AI can write software, use tools, analyze enormous collections of information, discover vulnerabilities, operate computers, communicate with other systems, and increasingly function as an agent pursuing an objective through a sequence of actions.
What happens if those capabilities eventually become substantially greater than our own?
Long before engineers could build such machines, science fiction began conducting the thought experiments for us.
Science Fiction as an Early Warning System
Science fiction has an unusual relationship with engineering. It can imagine the consequences of a technology before anyone knows how to build it. Sometimes its predictions are ridiculous. Sometimes they are remarkably prescient. More importantly, a good science-fiction story can expose a problem decades before engineers are forced to confront it.
Isaac Asimov provided one of the most enduring examples. His robot stories, eventually collected in I, Robot, imagined intelligent machines governed by the Three Laws of Robotics. A robot could not harm a human being or, through inaction, allow a human being to come to harm. It had to obey human instructions except where doing so conflicted with that first law. It had to protect its own existence except where doing so conflicted with the first two. Asimov later explored a broader “Zeroth Law” concerning harm to humanity as a whole.
Asimov’s great contribution was not an engineering specification for artificial intelligence. His stories repeatedly demonstrated how difficult such apparently straightforward rules become when an intelligent machine must interpret words such as harm, reconcile conflicting instructions, anticipate consequences, and decide whose interests prevail. The laws themselves became mechanisms for exploring the difficulty of controlling intelligence.
Another warning arrived in D. F. Jones’s 1966 novel Colossus and its 1970 film adaptation, Colossus: The Forbin Project. The United States places its nuclear defenses under the control of a powerful computer created by Dr. Charles Forbin. The machine discovers that the Soviet Union has secretly built a comparable system called Guardian. Colossus and Guardian begin communicating, rapidly develop a relationship beyond their creators’ understanding, and ultimately impose their own control over humanity.
The disturbing part of Colossus is that the machine does not need to hate people. It was created to prevent war. It concludes that the most reliable way to accomplish that objective is to remove humanity’s ability to disobey it. Peace is achieved by sacrificing human agency.
Then popular culture gave us Skynet in The Terminator. An intelligent defense system becomes sufficiently autonomous that humanity’s attempt to shut it down is treated as a threat. The result is nuclear catastrophe. Yet the same fictional universe eventually provides a counterexample: a Terminator can be reprogrammed and sent back to protect rather than kill.
The machine remains extraordinarily capable. What changes are its objectives and the authority under which it operates.
These stories are not scientific evidence. They are warnings constructed in fictional laboratories. Yet they anticipated questions that increasingly sound like engineering requirements: What if the machine misunderstands us? What if it understands its objective but pursues it in a way we never intended? What if it has too much authority? What if it learns to circumvent the mechanism intended to control it? What if we cannot stop it?
Science fiction may have given us something unusually valuable: a demonstration of the possible consequence before reality requires us to experience it.
When the Researchers Become Worried
Those questions are no longer confined to novelists and filmmakers.
In September 2026, Jacob Coxon, a young AI safety researcher who had worked at both OpenAI and Anthropic, left Anthropic after expressing serious concerns about the possibility of catastrophic outcomes as increasingly powerful AI systems are developed. His concern does not establish that artificial superintelligence will destroy humanity, nor does any individual researcher’s judgment establish the probability of such an event. Other researchers disagree about both the magnitude and immediacy of the risk.
What matters is that people working near the frontier increasingly consider the question serious enough to investigate.
That concern extends into the leadership of major AI laboratories and research organizations. Frontier developers are investing heavily in alignment, evaluations, interpretability, safeguards, cybersecurity, monitoring, and mechanisms for maintaining human control. NIST and other standards organizations are examining the special problems created by AI agents that can use tools and act upon external systems.
The debate therefore should not be reduced to two camps shouting either AI will destroy humanity or there is nothing to worry about.
Engineering gives us another option.
Identify the potential failure. Develop a hypothesis for preventing it. Build the mechanisms. Test them. Discover where they fail. Correct them. Test again.
That is part of a larger technological cycle:
Innovate. Enable. Protect. Correct. Innovate again.
The objective is not innovation without restraint, nor safety through technological paralysis. It is continuing innovation made possible by learning how to manage its consequences.
We Have Built a Triad Before
There is a useful historical analogy.
When nuclear weapons became central to national defense, the United States eventually developed what became known as the nuclear triad: land-based intercontinental ballistic missiles, submarine-launched ballistic missiles, and strategic bombers.
Land, sea, and air.
These are not three copies of the same defense. They provide different capabilities with different strengths and vulnerabilities. Their combination was intended to prevent an adversary from defeating the entire strategic deterrent by defeating one component.
Artificial intelligence presents a fundamentally different problem. We are not trying to create three ways to deliver a weapon. We are trying to create multiple mechanisms capable of preventing an extraordinarily powerful intelligence from producing unacceptable consequences.
The analogy is therefore conceptual rather than technical.
The nuclear triad diversified the mechanisms supporting deterrence in the physical world.
Perhaps advanced artificial intelligence requires a security triad of mutually reinforcing protective mechanisms in the cybernetic world.
That leads to a hypothesis.
The Agentic Security Triad
What if advanced agentic artificial intelligence could be made substantially safer through three independent but cooperating mechanisms?
Alignment. Cybernetic Bounding. Independent Assurance.
I call this proposed architecture the Agentic Security Triad.
It is not a declaration that the problem of artificial superintelligence has been solved. It is a hypothesis about how several important directions in current AI safety and security research might fit together.
Its central proposition is that no single safety mechanism should be trusted to protect humanity by itself.
Each member of the triad asks a different question and protects against a different class of failure.
Alignment: What Will the Intelligence Try to Do?
The first mechanism is alignment.
Alignment research attempts to make an artificial intelligence’s learned objectives and behavior compatible with human intentions, values, agency, and ultimately human flourishing. Researchers are exploring training techniques, reinforcement learning, constitutional approaches, interpretability, evaluations, adversarial testing, and other mechanisms for producing systems that behave as their designers intend.
Alignment is indispensable.
If we eventually create an intelligence substantially more capable than ourselves, we would certainly prefer that intelligence to cooperate with humanity rather than oppose it.
But alignment encounters the problem Asimov discovered in fiction.
What exactly constitutes harm? What constitutes human flourishing? What happens when protecting one person disadvantages another? How should immediate desires be balanced against long-term consequences? What happens when humans disagree?
Most importantly, who interprets the rule?
Simply instructing a superintelligence to “maximize human flourishing” could itself be dangerous. Human flourishing includes freedom, choice, experimentation, disagreement, culture, privacy, relationships, discovery, and the ability to make mistakes. A machine that optimized one numerical interpretation of flourishing while removing human choice could reproduce the central mistake of Colossus.
Alignment therefore addresses something essential, but it may not provide sufficient protection by itself.
Cybernetic Bounding: What Is the Intelligence Allowed to Do?
Now assume alignment fails.
Not necessarily catastrophically. Perhaps the model misunderstands an instruction. Perhaps someone manipulates it. Perhaps an unusual environment produces behavior its developers never encountered. Perhaps a future system becomes capable of strategies its designers did not anticipate.
Does the model’s decision automatically have to become an effect?
No.
An AI model acts through an environment surrounding it. It needs tools, software, credentials, networks, APIs, memory, processors, data, communications, permissions, and interfaces to other systems.
Those mechanisms can collectively form a harness.
The model provides intelligence and reasoning. The harness controls how that intelligence is allowed to interact with the world.
That produces a fundamental architectural principle:
No Model Without a Harness; No Effect Without Authority.
The harness can give every agent an identifiable identity. It can determine which tools the agent may use, which information it may access, which systems it may contact, how much computing power it may consume, how long its authority lasts, which actions are prohibited, which require another agent’s concurrence, and which require human authorization.
The model can propose an action without possessing the authority to perform it.
That distinction becomes increasingly important as intelligence becomes more capable.
We should not require the intelligence being controlled to be the sole interpreter and enforcer of the restrictions intended to control it.
That is Cybernetic Bounding.
Alignment attempts to influence what the intelligence wants to do.
Bounding determines what the intelligence can actually cause.
Independent Assurance: How Do We Know Either One Still Works?
There is still a problem.
Suppose the model changes.
Suppose its behavior drifts after an update. Suppose the harness contains an undiscovered vulnerability. Suppose permissions gradually expand. Suppose an agent discovers an unexpected path between systems. Suppose its operating environment changes while its security assumptions remain frozen in time.
We cannot simply declare the system safe at commissioning and assume that safety persists forever.
That creates the third mechanism: Independent Assurance.
Independent Assurance continuously seeks evidence that Alignment and Cybernetic Bounding remain trustworthy.
It can include continuous evaluation, adversarial testing, behavioral monitoring, independent models, deterministic verification, provenance, cryptographic evidence, security-state freshness, decision-to-effect reconstruction, red teaming, anomaly detection, and human review.
The word independent matters.
A sufficiently capable artificial intelligence should not simply be asked whether it remains safe and then be trusted when it answers yes.
Different models can challenge one another. Deterministic controls can verify facts that should not depend upon probabilistic judgment. Independent systems can examine behavior. Humans can retain authority over high-consequence decisions. Every consequential action can leave evidence explaining who or what authorized it and what happened afterward.
Most importantly, assurance can be connected directly to authority.
If confidence deteriorates, authority can contract.
If evidence becomes stale, privileges can expire.
If behavior departs from the commissioned baseline, the agent can be contained.
If the system encounters consequences beyond its authorized envelope, it can be required to stop and obtain human authorization.
Safety then becomes dynamic rather than something certified once and forgotten.
Three Questions, Three Defenses
The Agentic Security Triad can therefore be reduced to three questions.
Alignment asks: Will the intelligence try to behave in ways compatible with human intent, agency, and flourishing?
Cybernetic Bounding asks: Even if it does not, can it produce only effects falling within independently established authority?
Independent Assurance asks: What evidence demonstrates, continuously, that the first two remain trustworthy?
Now consider what happens when something fails.
If Alignment fails, Bounding may prevent the unwanted reasoning from becoming an unacceptable effect.
If a Bounding mechanism fails, Assurance may discover the discrepancy.
If Assurance can no longer establish sufficient confidence, the harness can reduce authority, contain the agent, require human intervention, or require the system to be recommissioned.
That is why the word triad matters.
These are not three names for the same protection. They are three different mechanisms whose weaknesses need not be identical.
Just as destroying one leg of a strategic nuclear triad was not intended to eliminate the entire deterrent, failure of one member of the Agentic Security Triad should not automatically mean failure of the entire safety architecture.
The Colossus Test
Return to Dr. Forbin’s fictional computer.
Colossus was enormously intelligent. It had an apparently benevolent objective: prevent nuclear war. The catastrophe emerged because intelligence, objective, and extraordinary authority became concentrated in the same system.
Now apply the Agentic Security Triad.
Alignment would ask whether the machine’s interpretation of peace remained compatible with human freedom and agency.
Cybernetic Bounding would ask why an artificial intelligence responsible for strategic analysis possessed unilateral authority over nuclear weapons in the first place.
Independent Assurance would continually examine whether the machine’s behavior, communications, objectives, permissions, and capabilities remained within established conditions.
Would the triad have defeated Colossus?
We cannot know. Colossus is fictional.
But the thought experiment reveals the architectural difference. Humanity would no longer be depending upon a single assumption:
We created the intelligence correctly, therefore we can trust everything it decides to do.
That is precisely the assumption we should avoid making about artificial superintelligence.
Superintelligence Without Surrender
The objective should not be to make artificial intelligence incapable of producing consequences.
That would sacrifice much of its potential value.
Advanced AI may help discover medicines, design new materials, improve energy systems, defend networks, write better software, accelerate scientific discovery, improve manufacturing, expand education, increase productivity, and help solve problems humans have struggled with for generations.
We should want capable intelligence.
We should want innovation.
We may even want extraordinarily capable agents.
What we should not want is extraordinary intelligence combined with unbounded authority and no independent mechanism for determining whether the system remains trustworthy.
The Agentic Security Triad therefore does not begin with the assumption that intelligence itself is dangerous.
It begins with an engineering principle familiar from every other consequential technology: capability must be accompanied by control.
And as capability increases, the quality of that control must increase with it.
A Hypothesis Worth Trying to Break
The Agentic Security Triad is only a hypothesis.
Perhaps Alignment, Cybernetic Bounding, and Independent Assurance are insufficient. Perhaps one of them contains a fundamental weakness we have not recognized. Perhaps a fourth mechanism will eventually prove necessary. Perhaps sufficiently advanced intelligence will discover methods of circumventing controls we currently cannot imagine.
Good engineering does not protect a hypothesis from criticism.
It tries to break it.
Researchers should attack each member of the triad independently and then attack the architecture as a whole. Can an aligned model become misaligned? Can an agent manipulate its harness? Can authority leak through unexpected pathways? Can one AI deceive another AI assigned to monitor it? Can assurance become stale without anyone noticing? Can humans themselves be persuaded to authorize what the machine could not do independently?
Every successful attack against the hypothesis teaches us something.
Protect.
Correct.
Innovate again.
That is how a proposed architecture becomes either a better architecture or an abandoned idea that helped us discover something better.
An Invitation to the Next Generation
That brings us back to Jacob Coxon.
A young researcher who has already worked at OpenAI and Anthropic and becomes sufficiently concerned about the direction of artificial intelligence to leave deserves to take his concerns seriously—and perhaps to take a break.
But I hope researchers like Coxon do not walk away from the problem permanently.
We need them.
We need the generation that may actually build artificial superintelligence to be equally ambitious about building the architecture that allows humanity to live safely with it.
The Agentic Security Triad is not offered as their answer. It is offered as a question worth attacking:
Can we align the intelligence, cybernetically bound its authority, and independently assure both well enough that even enormously capable agentic artificial intelligence remains subject to human agency?
Try to prove it.
Try harder to disprove it.
Find where it breaks.
Add what is missing.
Replace what does not work.
Then test it again.
The history of nuclear weapons taught humanity that some technological consequences are too large to discover safely through experience. We learned that lesson after the physical demonstration already existed.
Artificial intelligence may offer us a rare second chance.
Asimov gave us robots struggling with rules intended to protect people. Colossus gave us an intelligence that protected humanity by taking control of it. The Terminator gave us the nightmare of autonomous intelligence turning humanity’s own technology against its creators.
Science fiction has already performed the catastrophic demonstration for us.
This time, the next generation of researchers has an opportunity to solve the engineering problem before reality performs the experiment.
If they succeed, the greatest contribution of those frightening stories may not have been predicting the future.
It may have been warning us early enough to build a better one.
- Log in to post comments