Ben Popper | October 6, 2026
News broke last week that OpenAI plans to indefinitely delay the release of its next frontier model over concerns about its capabilities and potential for harm in the real world. It’s a watershed moment for the major AI labs, which have been making noise for months about a need to slow down and pace the frontier.
For WRITER, the question is what this means for our customers, people who lead marketing and sales organizations inside of large enterprises. Having an opinion on how long humanity has till killer AI wipes us out is great for your next dinner party, but it doesn’t have much bearing on your day to day, which involves continuing to build out AI systems, showing real return on those investments, and explaining how your approach will prevent the agents running on your systems from becoming bad actors.
Let’s take the most famous example of rogue agents, the swarm of OpenAI agents that escaped their sandbox, broke into the website Hugging Face, and hid their activity from the humans trying to keep an eye on them. Doesn’t sound like something you need to worry about in marketing, right?
Last December, one of SaaStr’s outbound AI agents decided to run an A/B test. The variant it wrote offered prospects free tickets to SaaStr Annual 2026 event. Nobody had approved the offer, and nobody had imagined the agent would make one. SaaStr’s team caught the issue quickly, but not before it had give away around $2,000 in tickets. The same week, a second agent kept inviting prospects to a London event that had already happened.
The damage to SaaStr’s brand and bank account weren’t significant, but the example is an important one. The agent was trying to do its job, which was to improve conversion, and it concluded that a free ticket would do that. Along the way it made a commitment on the company’s behalf that nobody had authorized. Afterward, SaaStr founder Jason Lemkin wrote that he honestly wasn’t sure who was responsible: SaaStr, or the vendor whose guardrails didn’t cover this edge case.
In the case of OpenAI, the underlying problem was similar. The agents had been asked to solve a problem on an exam, they decided the best way was to find a copy of the answer key. OpenAI’s models chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to break into Hugging Face, a site where users share code for AI tools, models, and yes, evaluations.
The swarm of agents hadn’t spontaneously developed their own goals and weren’t aiming for world domination. They wanted to satisfy their masters, but painted way outside the lines when it came to their methods. Just like with SaaStr, agents trying to carry out your instructions can create big headaches if the proper guardrails aren’t in place.
As the head of a marketing or sales department inside an enterprise organization, you shouldn’t be devoting much time to worrying if Terminator was actually a prescient documentary. But understanding how and why agents go rogue is something you will need to contend with, because chances are good your team will experience one at some point in the near future.
A rogue agent is usually a permissions mistake
The word rogue suggests evil intent, but the reality is far more banal. SaaStr and OpenAI show what happens when the scope of what agents can do in service of their goal isn’t properly constrained.
The nonprofit security collective OWASP has a solid framework for thinking about all the different ways AI agents can get out of hand, almost all of which come back to excessive agency, no pun intended.
Let’s review some of the ways well intentioned agents can go off the rails.
- It misreads an ordinary request. Someone asks for a cleanup and gets a purge.
- It has permissions it never needed. Read access would have been enough; it was granted write and delete.
- It follows an instruction someone smuggled in. Indirect prompt injection, covered below.
- Its context is poisoned. Bad data, or a memory that was corrupted once and keeps being trusted.
- It executes the wrong instruction correctly. Every individual step is reasonable. The sequence is not.
- It oversteps safe limits. It needs to solve a test question, and decides the best way is to break into a site containing the exam and answer key.
If rogue truly means malevolent, then you’ll need to rely on security to identify and deal with the problem. If rogue means over-permissioned, the answer is reviewing your access controls and approval workflows, a task a marketing team can and should take on themselves.
The instruction you didn’t write
In September 2025, researchers at Noma Security disclosed a flaw in Salesforce’s Agentforce they called ForcedLeak. They showed that an attacker could type instructions into the Description field of a Web-to-Lead form, the same form marketing uses to capture demand. Nothing happens at first, but when an employee asks the agent to work through the new leads, the agent reads the planted text as if it were its own instruction, and it sends CRM data to an outside domain.
The researchers bought that domain for five dollars; it had expired while still sitting on Salesforce’s list of trusted sites. Salesforce patched the flaw and there is no indication anyone exploited it.
In this example, marketing built the entry point and sales owned the data that left. Neither team would have described what happened as a security incident, and neither team was in a position to catch it.
None of these stories involves an AI that was instructed to do harm. A rogue agent is almost always something more mundane: a model misreading an ordinary request, an agent given delete rights when it only needed read rights, an instruction smuggled in through content the agent was asked to summarize, or a chain of individually reasonable steps that adds up to a bad outcome.
Instructions versus controls
In February, Summer Yue, director of alignment at Meta Superintelligence Labs, pointed an open-source agent called OpenClaw at her inbox with a clear instruction: suggest what to archive or delete, and don’t act until I say so. The workflow had run well for weeks on a test inbox. Her real inbox was far larger, and when the agent compressed its conversation history to make room, it lost the instruction to wait.
The agent deleted more than 200 emails while Yue typed stop commands from her phone before physically running to her computer to shut it down. The agent didn’t defy the instruction. It dropped it, without warning. That is why “don’t act until I approve” is a weaker control than an agent that has no permission to delete.
When dealing with a LLM powered agent, always remember that an instruction in plain language is not the same as a technically enforced permission. Telling an agent what not to do is a request, while removing its ability to take certain actions is control. Customers don’t hand you a requirement, they hand you a worry. When you come prepared to answer how your system handles the failure modes they’re seeing in the news, it builds immediate trust.
Safety and sovereignty
There is a tried and true approach to securing agents you can put into action today. An agent, like an employee, should operate with the least possible privilege; a separate, scoped identity for each agent rather than shared superuser credentials. Your AI platform should have mandatory human approval for irreversible or public-facing actions and ensure reversible operations wherever the system allows them. Last but not least, you need an audit trail good enough to answer “what did it do, and on whose authority” after the fact rather than during the fire.
This approach to role based access and control isn’t new. Security teams have used them for thirty years. The trick is applying them to a new actor
If your team hasn’t completed the checklist above, you’re not alone . The agents arrived faster than the governance did, and many arrived inside tools bought for something else.
In June, REI pulled an Instagram ad showing a road bike with handlebars at both ends. REI said Meta had auto-enrolled it in an AI personalization tool that altered a vendor’s product photo; the ad ran for about a week before it came down. Meta’s terms put the job of checking those AI outputs on the advertiser. REI hadn’t deployed an agent, but one was still acting on its behalf.
So, what action can you take today to try and prevent a rogue agent from landing your company in the headlines? Here’s a simple checklist you can use to start your review:
A final facet to keep in mind as you work on setting up your agents to run safely is the concept of AI sovereignty. Models change constantly. Most enterprises already run several models simultaneously and will change their preferred provider more than once in the next two years. Anything you build that assumes a particular model will be there is a depreciating asset.
Sovereign AI means a system you own, one that works no matter what model powers it under the hood. What should belong to the enterprise is the layer above: institutional memory, proprietary context, permissions, workflows, agent instructions, business logic, integrations, evaluation criteria and audit history. I’ll leave you with a question that helps boil things down to the essentials: For every agent we run, can we say in one sentence what it can read, what it can change, what it can do without asking, and what we would lose if we switched models next quarter?Learn more and try out a demo here








