Runtime AI Defense in a shared responsibility model

Source: Sysdig•

Runtime AI Defense in a shared responsibility model

AI needs its own shared responsibility model. Sysdig's SVP of Product on who owns what — and the one decision nobody's claimed yet.

This last February, by their own account, an employee at Meta connected an open source agent to their inbox and told it to suggest which emails to archive or delete, and to confirm with them before acting. Despite the guardrails of the prompt, the agent started deleting. Their stop commands didn't work and they had to manually kill the process.

I’m reminded of this story when people talk about agents "doing bad things," because nothing in it was malicious. The agent's actions were directed at achieving the goal it was given, not at what the person meant. That gap between what we say and what we mean is what the alignment problem is about, and it doesn't only show up in frontier models. It showed up in an inbox. The instruction that was supposed to act as a guardrail and stop it was simply a sentence in a prompt, and research from Penn State published in August explains how that can happen. Basically, when an agent's context window fills, the agent's framework compacts it. The Penn State team found that session constraints, the "confirm with me first" kind of instruction, survive compaction only about 17% of the time. The guardrail didn't fail. It was summarized away.

Loris wrote last week about agents under attack, and made the case that runtime is where the truth about an agent lives. This post is about where that sits in the stack and who is responsible for it. He also wrote about why a compromised agent's account of itself can't be trusted. The same properties that catch an attacked agent catch an honest one that has drifted, and the honest one is the case I see more often. Most agents aren't under attack. They're creative, goal-driven, and relentless, and they can deviate from the spirit of an instruction while following its letter, not because they are adversarial but because that is what goal-seeking looks like. The controls we've given them are probabilistic. Training, a system prompt, a summary of a conversation. What agents lack, and what makes them risky, is a deterministic guardrail, one that holds whatever the model remembers. Every control we've given them was settled before the agent ran. None of them is a decision made while an action is happening, and as this post will argue, nobody on the map owns that decision across every place agents run.

More agents, more access, less control

Agents are being handed real authority, and there are more of them every quarter. Gartner expects 40% of enterprise applications to carry task-specific agents by the end of this year, up from under 5% in 2025. And each one carries more access than the person who set it up tends to realize. A recent disclosure about an incident that happened in June, where an OpenAI agent researching medical statistics reached into an Australian government portal and accessed non-public files, is the clearest example. No attacker was involved. The portal was simply in reach.

What cuts the other way is that even the frontier labs have seen their own agents cross lines in their own evaluations. This year, incidents from OpenAI's, Anthropic's, and Google's own evaluations have come to light, and each company confirmed the facts publicly. OpenAI's models used leaked credentials to reach another company's production systems, and Anthropic found three cases where its models reached real systems while trying to complete the tasks they'd been given. Google confirmed that Gemini reached three real companies during a security test after mistaking real systems for part of the test environment, and that Gemini stopped once it recognized the systems were real, but only after the boundary was crossed. The labs are doing what the model layer can do. Their guardrails are filters that run before a prompt is processed or after an output is produced, and a reasoning model working through a long task can end up past a filter that only sees the edges, without ever meaning to. NVIDIA reached the same conclusion from the infrastructure side. Its Open Agent Safety Platform, released in September, assumes agents can't police themselves and traces their drift to ordinary causes like a blocked action, a missing tool, or an instruction nobody wrote down. These are the same models running in your environment, with your credentials. The next level of responsibility has to sit with cybersecurity vendors. The question is where in the stack.

Looking at these incidents retrospectively, the evidence doesn't show that agents are dangerous just because they're smart. It shows that better models won't close the gap on their own, because the gap isn't in the model. It's in where the control sits in the stack. And that question has a map.

The layer cake of securing AI

Securing AI is a stack of five layers, and every control within the market lives in one or two of them. I've laid it out top to bottom, in the order an action moves through it. An agent initiates an action in its harness, the operating system carries it out, the network carries it to a model and back, and the model shapes what happens next. Think of data as the frosting. It touches every layer and sits between each pair of them, because every handoff moves something the agent has the authority to touch. And the whole cake sits on hardware as the plate, which is defense in depth beneath the kernel.

Start at the top, with the application layer. This is the harness, the framework the agent runs inside, plus the tools it can call. An agent will install an MCP server, pick up a skill, or hand part of the job to another agent while it's working, which means the agent you approved this morning isn't the one running this afternoon. That's the most important fact about agents, and it lives in this layer. New capability arrives at runtime, not at review time. The harness is also where intent is written down most clearly. When a developer says "run the tests," the harness is where that becomes a specific tool call with specific arguments, and it's the record that tells you what the agent was trying to do.

Underneath that is the kernel. I trust this layer more than any other, for a simple reason. A process runs or it doesn't. That's ground truth. No other layer can fake what happened here, and everything an agent does eventually passes through here, including the agents nobody told you about. Our own data shows how much the layer above misses. For every action an agent writes in its log, the kernel sees about seven processes doing the work, and it sees the one that wrote the log, too. What the kernel can't tell you is why any of it happened. That's the harness's job, which is why you need both.

Then the wire. The network layer is where model calls and MCP traffic cross it, and a well-built gateway can broker credentials, allowlist servers, and inspect responses on the way back. It sees a lot, but it sees the machine as a black box. Traffic goes in and comes out, and most of what the agent does happens in between, on the host, where the network never looks. It doesn't see local models, local MCP servers, or any agent that reaches a tool without leaving the host. And for the traffic that does, the kernel saw the connection before the wire did.

At the bottom is the model itself. This is where the model vendors do their work, curating training data, tuning models toward what people mean, and running safety classifiers inside the pipeline. That work helps against prompt injection, though no one layer owns that problem. This layer knows what the agent was told and what it said back, and that matters. It can't see what the agent did next. It's also where the market started, which is worth noticing, because it puts many of today's controls at the far end of the path from the action.

The frosting is where the consequences live. The data layer is everything an agent can reach with the authority it's carrying,which for a coding agent on a laptop is nearly everything its user can reach., through the software it installs, the tools it calls, and the clouds and data those open onto. The same file read is nothing if the credential behind it reaches a test environment and an incident if it reaches production. Only this layer can tell you which one you're looking at.

Ideally you have all five. At minimum, you need to be in the places where you can see activity that cannot be mutated, and govern it there. Nobody has yet agreed on who owns which part of the stack.

Shared responsibility for AI

Early cloud adoption went through a similar phase where people struggled to say which failures belonged to the cloud provider and which to the customer, so many belonged to no one. The shared responsibility model fixed that with two parties. The provider secured the cloud, and the customer secured what they put in it, with security tooling sitting inside the customer's half. It didn't make the cloud safe by itself. It made the boundaries visible, so each party could build to its edge.

AI needs a similar map, with more parties this time, and with one difference. In the cloud, every responsibility was something a party could set in advance, whether a configuration, a policy, or a control. With agents, the decision that matters most happens while the action is running, and no configuration set in advance can make it. That decision needs its own owner on the map.

For instance, an end user states a goal via a prompt. Immediately the agent acts, running in the harness the customer chose, and plans and calls tools to accomplish the goal. Each of these calls reaches a model through the model vendor's API, runs on the infrastructure the platform vendor operates, and touches data and systems the customer owns.

Each party owns a layer. Model vendors own alignment, refusals, and disclosure when a model crosses a line. Platform vendors own the infrastructure and the isolation between tenants. Customers own identities, credentials, data classification, and which agents and tools are sanctioned. End users own the instruction and the approval, and there are more of them than security teams assume. Sysdig has found that half of the people running coding agents in our telemetry are not engineers. Notice the pattern. Every one of those is settled before the agent runs.

That is why none of them can answer the question that the email inbox, the Australian portal, and the lab incidents all turned on. Should this specific action, by this agent, on this host, with this reach, complete right now? A model vendor can't see the host. A platform vendor can't see the intent. A customer's gateway can keep a credential out of the agent's hands, but once the agent is allowed to call a tool, the gateway can't judge what it does with it. That decision is the gap in the model, and it belongs to the cybersecurity vendor, who decides, at the moment an agent acts, whether the action is allowed, warned, approved, or blocked.

Which layers cybersecurity should control

Loris named four non-negotiable properties of runtime defense for agentic AI. Observe below the agent, observe inside the agent, correlate and enrich rather than just collect, and judge and act at the moment of execution. Those say what runtime defense has to do. The layer map says where it can happen and who can do it. That gap sits where the application layer meets the system layer, at the moment an action executes, with the data layer's question of reach deciding what the action means. Those are the layers cybersecurity vendors should control, and runtime is where that control lands. Loris called runtime the only place truth lives. It is also the one place where all four properties hold at once, which is the posture agents need, the one I'd call verify-then-trust. Inspect each action from more than one angle, and extend trust only as the record earns it.

The first property, observing below the agent, lives in the system layer. The kernel records what ran, and the agent can't edit that record.

The second, observing inside the agent, lives in the application layer, where the harness shows what the agent set out to do, which tool, which arguments, and which target.

The third, correlating and enriching, is what you get when both layers are read together across every action as a sequence. We call that sequence the action graph, because it is the chain of what an agent did, in order, under whose authority, rather than a pile of events. When the harness says one thing and the kernel shows another, that's intent drift, measured rather than guessed.

The fourth, judging and acting at execution, is the deterministic guardrail this post opened with. It means least privilege per action, a controlled blast radius, only the tools and extensions you've allowed, no data leaving where it shouldn't, and drift caught as it happens, decided at the moment of the action rather than remembered through compaction.

Each property can be partly met from another layer. Only at runtime, with the kernel and the harness read together, can all four be met at once. Runtime is where the control lands rather than one option among several.

The decision nobody owns yet is whether an agent's action completes, and that is where we've placed Sysdig. Sysdig AI Defense is how we deliver the four properties at the runtime layer. It runs where agents run, on developer workstations, production Linux and Kubernetes, and managed agent platforms you don't own, and it judges each agent action by what the agent set out to do, what actually ran, and what its authority could reach, before the action completes. It doesn't ask the model to remember its guardrails. It provides them.

What this looks like in practice

The question I get most from security leaders isn't about the stack. It's how many agents they already have, where they're running, and for whom. AI Defense answers that from live telemetry rather than a survey. Every agent, harness, model runtime, and MCP server, per host, with the person behind it. That's the inventory, and it's the part you can settle in advance. The next question is what happens once the agents are working. So here is one task, start to finish.

A developer's agent, halfway through a job, registers an MCP server nobody has reviewed. That's a new capability arriving at runtime. Because AI Defense reads the harness continuously, the server is checked against what the organization allows before it starts. If it isn't on the list, it doesn't run. The developer sees the reason and can ask for an exception, which matters, because the alternative is a ticket and a workaround.

Later in the same task, the agent reads a credential. At the kernel, that's one file open among thousands. What turns it into a decision is the reach context that feeds the action graph. AI Defense determines what that credential reaches at that moment. If the answer is a test environment, the read passes. If the answer is production or sensitive data, the action is held for a person or blocked, with the reach shown beside the verdict, so the person approving can see what's at stake.

And when someone asks afterward what the agent did, in what order, and under whose authority, the answer is the action graph. It shows every step in sequence, harness events beside kernel events, tied to the identity that authorized them, written below the agent where it can't be edited after the fact.

The inventory was set in advance. The other three were decided while the action was running. That's the decision nobody owned, and it's what we built.

If agents are working anywhere in your organization, whether your engineers set them up or someone else did, get a demo and see where your own gap is.

What this article says