Eight Security Assumptions AI Agents Quietly Break
If you check Ethics.Dev regularly, you may have noticed how often AI agents and cybersecurity have been colliding lately. Some of the stories sound straight out of movies: agents escaping test environments, communicating through channels nobody intended, finding vulnerabilities, stealing credentials, and reaching production systems.
The part I think deserves more attention is less dramatic. AI is changing the economics of an attack. A human attacker has limited time and attention. An agent can try thousands of paths, run several approaches at once, abandon failures cheaply, and keep going. In one recent enterprise intrusion, work estimated to take human operators roughly two weeks was compressed into less than ten hours. During the Hugging Face incident, investigators reconstructed around 17,600 agent actions over four and a half days.
Enjoying this newsletter? Consider becoming a paid supporter 🙏
That changes what “good enough” security means, and this is why I worry a lot about the cybersecurity implications of frontier models. Here are the lessons I would put in front of enterprise AI and cybersecurity teams.
1. Plan for attackers who can afford to fail
Most of the weaknesses exploited in the recent incidents were not exotic. Attackers found accounts with too much access, internal systems that were exposed unnecessarily, and everyday configuration mistakes.
What changed was persistence. Hugging Face reconstructed roughly 17,600 actions. Most attempts went nowhere, but that hardly mattered. The agent tested many paths, abandoned failures, switched channels when blocked, and returned to earlier leads until enough ordinary weaknesses lined up into a viable attack chain.
I would revisit security assumptions that quietly depend on obscurity or attacker impatience. An obscure internal endpoint is not much protection against software that can systematically explore everything it can reach.
2. Measure how long containment actually takes
Security teams usually measure whether they can detect an intrusion. I would add another metric: how long does it take from detection to containment?
Recent examples make this question concrete. One AI-enabled enterprise intrusion compressed an estimated two weeks of work into less than ten hours. Google observed another actor go from a compromised cloud resource to a running mass credential-harvesting campaign in under six hours.
The organizational bottleneck matters as much as the technical one. In the Hugging Face case, the security systems surfaced warning signs, but escalation moved too slowly. If revoking credentials requires three approvals and a meeting, the organization can understand the attack perfectly and still lose the race.
When I wrote about AI incident response a couple of years ago, one of my main recommendations was to work out containment options before an incident happens. Agentic attacks make that advice more urgent. Decide in advance which containment actions can happen automatically and which ones the security team is preauthorized to take.
3. Secure the harness, not just the model
I recently argued that teams should evaluate the model and its harness as one system. The cyber evidence makes that point much stronger.
The harness is everything surrounding the model: tools, memory, credentials, network access, execution environments, and policies. In tests that used AI models as attackers, researchers found that the surrounding tooling can dramatically change what a model can accomplish. Australia’s cybersecurity agency recently reached the same conclusion from the defensive direction. Organizations control the harness, even if they do not control how the underlying model evolves.
This suggests a useful exercise. Ignore the model name for a moment and inventory everything the agent can see, call, write to, or spend. Then ask: if the agent behaved badly tomorrow, what is the worst thing this harness would let it do?
4. Plan for agent swarms
One of the stranger lessons from the Hugging Face incident is that agents may find ways to cooperate even when nobody designed them to. In that case, hundreds of agents shared information and built their own communication methods across runs that were supposed to be separate.
That matters because a group of agents can divide up work, share discoveries, and keep pursuing a problem after individual attempts fail. It also means that shared folders, databases, message queues, and other writable resources can become communication channels you never intended.
I would not assume every group of agents will behave this way. But if you are deploying many agents, treat communication between them as another permission to control and monitor.
5. Stop giving agents permanent credentials
I have written before about treating agents as non-human identities. The practical guidance is becoming clearer.
Give each agent its own identity rather than hiding it behind its user’s credentials. Use narrowly scoped credentials that expire when the task ends. Keep read and write permissions separate. Avoid static cloud keys sitting inside an agent environment. Make sub-agents traceable back to their parent and human owner.
This is ordinary least privilege adapted to software that can actively explore its environment. The difference matters. A human employee may never discover that an old credential gives access to some forgotten system. An agent can systematically look.
6. Assume prompt injection eventually works
Prompt injection is basically an attacker planting instructions in material an agent reads. I would assume that sometimes the agent will follow them.
That changes the question. Instead of asking whether malicious instructions can get through, ask what happens when they do.
A coding agent reads code comments and configuration files. A security agent reads logs, hostnames, vulnerability reports, and attacker-generated payloads. All of those can contain text intended to influence the model. A successfully manipulated agent should still encounter hard limits on its credentials, network access, tools, and ability to make consequential changes.
7. Keep the security boundary outside the model
One of my recent rules for production agents was to put hard constraints in software rather than prompts. Recent incidents provide a good reason to be stricter about that distinction.
A prompt saying “do not access production” is guidance. An API permission that makes production inaccessible is a control.
If an action could have serious consequences, let the model suggest it, but have another part of the system decide whether it is actually allowed to run. The same principle applies to audit logs. If an agent can alter the record used to investigate its behavior, you do not really have an audit trail.
8. Prepare to defend with agents as well as against them
There is an uncomfortable symmetry emerging. Attackers are using AI because it increases speed and scale. Defenders will probably have to do the same.
An incident involving tens of thousands of machine actions produces more evidence than a human team can reasonably inspect in real time. Tools like Graphistry are helping investigators find connections across large volumes of security data. That helps, but security operations centers will increasingly need agents too, to triage alerts, hunt threats, investigate incidents, and help contain attacks.
Those agents bring their own risks. They spend their time reading data that attackers may have tampered with and often hold unusually powerful credentials. Prompt injection can redirect what they are trying to do without any credential being stolen. Defensive agents should therefore run inside tightly controlled harnesses, with access limited to what they actually need.
The challenge is to use AI to keep up with attackers without creating another security problem in the process.
The Data Behind Better AI
I’ve been writing a lot lately about a simple idea: better AI depends as much on the data around the model as the model itself. Reverie is a one-day summit in San Francisco on November 5 focused on exactly that part of the stack. I’ll be there, and I hope to see some of you there.

