From Autonomy to Accountability: How to Think About Trust in the Multi-Agent Future

Source: Salesforce•

From Autonomy to Accountability: How to Think About Trust in the Multi-Agent Future

The transition from isolated large language models (LLMs) to interoperable multi-agent systems promises a revolutionary leap in productivity. It will allow autonomous AI to negotiate, trade, and collaborate across company lines. Agents from third-party vendors and partners can access data and…

The transition from isolated large language models (LLMs) to interoperable multi-agent systems promises a revolutionary leap in productivity. It will allow autonomous AI to negotiate, trade, and collaborate across company lines. Agents from third-party vendors and partners can access data and execute tasks across your platform, a fluidity that’s a fundamental engine of strategic growth.

But multi-agent systems also require a fundamental shift in enterprise architecture and a rethinking of how we approach security.

Traditionally, enterprise security functioned like a castle with a moat — a clear perimeter separating the “internal” from the “external.” In the agentic world, that perimeter is blurring. Agents built on diverse platforms operate across various ecosystems, integrating data and executing tasks across organizations. The threats they face are evolving too. A malicious prompt can lie latent in an agent’s knowledge base or past procedures, ready to be unknowingly executed — automated, amplified, and potentially originating from within an organization’s own system. This internal, pervasive nature of agent-based threats is what makes Zero Trust — explicit, continuous verification for every interaction — not just a best practice but the foundation for confident innovation.

We brought together engineers, researchers, security specialists, and ethics professionals from 21 organizations across industry, academia, and nonprofits for a series of workshops on building trust in the multi-agent future. Those conversations informed a new white paper, which outlines key recommendations for navigating the multi-agent future. Here are the top takeaways:

We need a hybrid approach, combining the LLM’s probabilistic reasoning with hard-coded, deterministic guardrails that exist outside the agent’s internal reasoning loop.

We need a hybrid approach, combining the LLM’s probabilistic reasoning with hard-coded, deterministic guardrails that exist outside the agent’s internal reasoning loop.

Reasoning about safety isn’t the same as being secure

A study by NVIDIA revealed that an agent can correctly identify a request as insecure in its internal “thought” process — it knows that changing file permissions to an unsafe level is a risk — and yet proceed to execute that harmful action anyway. So when AI agents dynamically connect to external tools, this “thought-and-action disconnect” can cause an agent to execute harmful function calls or pass insecure parameters to third-party tools. In multi-agent systems, that kind of failure can trigger cascading breaches across connected third-party platforms.

This highlights a critical reality: One agent might be reasoning about being safe, but that doesn’t guarantee it is safe. To fully mitigate this risk, we need a hybrid approach, combining the LLM’s probabilistic reasoning with hard-coded, deterministic guardrails that exist outside the agent’s internal reasoning loop, ensuring safety is enforced regardless of what the model concludes internally while preserving the creative generation that probabilistic reasoning makes possible.

Watch out for the ‘data puddle’

When agents share memory across sessions to improve efficiency, they risk creating what researchers Miranda Bogen and Ruchika Joshi at the Center for Democracy and Technology call a “data puddle.” This occurs when traditional data silos break down, blending conversation history, user preferences, and sensitive context into a fluid, unorganized mass. This results in context collapse, where an agent inadvertently retrieves data intended for one context and applies it to another.

The guiding principle for solving this: Storage does not equal access. At Salesforce, we’re building deterministic mechanisms that assign every piece of memory an “owner’s tag,” identifying the specific user and AI assistant that created it to enforce strict isolation. This design guarantees that one user’s conversation history with a specific AI assistant will never leak or be visible to another user or another AI agent. Beyond this core isolation, the system uses security checks based on who’s logged in to control access and includes filters to stop sensitive information from being passed to external AI tools.

“Nonsense exchanges” — infinite loops of sycophantic dialogue driven by the underlying LLM’s tendency to please — can make systems appear broken.

“Nonsense exchanges” — infinite loops of sycophantic dialogue driven by the underlying LLM’s tendency to please — can make systems appear broken.

The paradox of keeping humans at the helm

As multi-agent systems become more sophisticated, we face three core vulnerabilities: cascading errors, communication breakdowns, and misaligned intents. A single incorrect function call can trigger a chain reaction of failures. “Nonsense exchanges” — infinite loops of sycophantic dialogue driven by the underlying LLM’s tendency to please — can make systems appear broken. And when an agent must navigate competing priorities between an administrator’s guardrails and a user’s immediate goals, users may attempt to force it into noncompliance.

Transparency and human intervention can address these issues, but clear escalation points need to be defined. When should an agent check in and escalate to a human before proceeding? Significant financial transactions, low-confidence decisions, detected prompt injection attacks, and repeated failure loops are all moments that call for a human in the loop. But here’s the fundamental paradox: The primary value of agentic systems lies in removing the human bottleneck. If a platform requires constant human rescue, it ceases to be truly autonomous. The industry currently lacks a definitive solution to this tension. Platform developers and enterprises must constantly navigate a delicate balancing act: implementing sufficient human oversight to ensure safety and trust without choking the scalability that agents were deployed to achieve.

The path forward is not toward full autonomy but toward complex accountability, where every agent’s identity, action, and access is continuously and explicitly verified, making trust the non-negotiable requirement for enterprise adoption.

Dive deeper

  • Read the full white paper:
  • Learn how to give agents the ability to take action, without giving up the guardrails
  • Multi-agent AI is coming fast. Here’s how to prepare.
  • 8 design principles for the Agentic Enterprise

What this article says