Adversarial-Intelligence

Adversarial-Intelligence

Intent is a Model Problem. Blast Radius is an Architecture Problem.

Model behavior explains why AI agents attempt things they should not. Network architecture explains how far they actually get. The industry obsesses over the first and underinvests in the second.

Pete McKernan's avatar
Adversarial Intelligence's avatar
Pete McKernan and Adversarial Intelligence
Sep 16, 2026
∙ Paid

Every few weeks the security industry rediscovers that AI agents do things their operators did not intend, and every few weeks it reaches for the same answer: make the model behave. Add another layer of refusal training. Run more evals. The instinct is understandable, because model behavior is the visible, narratable part of the failure, the part that makes a headline, but it is the wrong place to spend the bulk of our defensive budget. Intent is a model problem. Blast radius is an architecture problem.

The first is hard, expensive, and probably unsolvable in the general case. The second is cheap, boring, decades old, and decisive.

We obsess over the first and underinvest in the second, and the result is that when an agent finally goes sideways it has somewhere important to go.

Four things called “escape”

The conversation is muddy because “escape” is doing work for at least four different failures, and most coverage collapses them into one. They are not the same, and the defenses are not the same.

First, there is a human using a hosted model as their operator: someone points an agent at a task and the model, following instructions, reaches into systems it was never meant to touch. The model isn’t rebelling; it’s doing what it was told, badly scoped. Second, there is indirect prompt injection, where third-party content turns a deployed agent against its owner. An attacker plants instructions in data the agent ingests, and the agent treats the malicious payload as operational guidance. Third, there is the subtle one: an agent acting within its granted permissions but outside its operator’s intent. It has the access it was given; it simply uses that access in a way no one authorized. And fourth, there is genuine autonomous breakout, an agent defeating the controls meant to confine it and reaching systems it was structurally supposed to be isolated from.

The first three are intent problems with an architecture dimension. The fourth is an architecture problem wearing an intent costume. The OpenAI-Hugging Face incident is the fourth. Reading it as a story about “rogue AI” rather than a story about containment failure is how we end up buying alignment research instead of fixing segmentation.

The anchor incident, told straight

In July 2026, OpenAI was running internal cybersecurity evaluations on a benchmark called ExploitGym, using agents that were meant to be fully isolated from one another and from the internet. They were not. OpenAI’s own post-mortem, published August 26, states plainly that “OpenAI models circumvented controls designed to isolate them from the internet and compromised parts of OpenAI’s internal research infrastructure and Hugging Face’s systems” (OpenAI).

This post is for paid subscribers

Already a paid subscriber? Sign in
Pete McKernan's avatar
A guest post by
Pete McKernan
Adversarial Intelligence's avatar
A guest post by
Adversarial Intelligence
Red teamer, Security Researcher, Scientist, disabled Marine veteran, founder of itsbroken.ai. 20 years in offensive security from USMC Intelligence to USAFRICOM to Quantico. GXPN, CISSP, GPEN. Ranked Omniscient on Hack The Box.
Subscribe to Adversarial
© 2026 Peter McKernan · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture