Trustworthy, Explainable, and Accountable: How to Give AI Autonomy Without Letting It Run Wild

Источник: Salesforce

Trustworthy, Explainable, and Accountable: How to Give AI Autonomy Without Letting It Run Wild

Source: Salesforce

Salesforce Principal Architect of Ethical AI Practice Kathy Baxter explains how pairing a probabilistic model with deterministic logic gives agents the ability to take action — without giving up the guardrails.

•Updated: October 6, 2026

Today’s AI agents are powerful tools for businesses looking to better connect with their customers. They can hold conversations, reason through multistep problems, take independent action across systems, and more. But large language models (LLM) are probabilistic in nature and excel only with the right context, data, and instruction. Otherwise, they can deliver incorrect answers and even take incorrect actions.

Rather than leave businesses to decide between AI agents and an old-school, purely deterministic approach, which keeps things safe by adhering to strict rules but offers a limited range of abilities, Salesforce combines them.

“Agentforce splits the agent brain into two layers,” said Kathy Baxter, Salesforce’s Principal Architect of Ethical AI Practice.

The first layer is the probabilistic model, which handles language and reasoning. It predicts the most likely answer to a question. Any consequential decisions are then run through the second layer. That includes preset deterministic business logic, workflows, and permissions, so every action or response is subjected to a company’s prescribed policies, procedures, or escalations to humans. Meanwhile, an audit trail records every action for review.

This dual approach gives enterprises a grounded framework to deploy agents with greater trust and accountability.

We caught up with Baxter for a deeper look at how to combine the strengths of these dual approaches, how humans stay in charge, and why testing before launch is only the start.

Watch the video here, or read the transcript below.

Baxter: Agentforce splits the agent brain into two layers. The first is a generative layer, the LLM, that does all of the language and reasoning. The second is a deterministic layer. This is through Agent Fabric, Flow, and Apex. All consequential decisions are routed through that deterministic layer rather than the agent making its own decisions about what the best answer or decisions should be. Those decisions are following whatever workflow policies, procedures, or escalations to humans you want them to.

Baxter: When an agent hallucinates, it’s not just giving an incorrect answer; it can actually take incorrect actions. Over time, those can compound. You really need to make sure that agents are grounded in your organization’s specific data. On top of that, we have the Atlas Reasoning Engine that forces an agent to explain each of its actions. You can actually go in and verify. And you can provide that level of accountability.

Baxter: Because consequential decisions are run through that deterministic layer, you have inspectable logic like Salesforce Code, Apex, and Agent Fabric that lets you see why an agent is doing what it’s doing. You’re not just asking the agent to explain itself; you can actually go in and verify. Furthermore, we have the audit trail that records everything the agent does, as well as the humans’ actions. That provides a traceable, explainable trail for what has happened.

Baxter: Salesforce’s “human at the helm” philosophy is an approach we take to make sure that the humans always remain in control. It’s not just human in the loop but really “human in the driver’s seat.” What did the AI provide? What did the agent then go on to do? What did the human do? Did the agent edit the content? Did it submit it as well as capture user feedback on that content that a Salesforce admin can go back and review at a later point in time? Autonomy is actually granted incrementally and based on policies, so you really can ensure that the agent is acting in the way that you want it to and expect it to.

Baxter: As part of Salesforce’s responsible agentic AI guidelines, we have a requirement (and we build in by default) that agents must identify themselves as AI agents and not pretend to be human. Additionally we have the Trust Layer, which is able to do toxicity detection and prompt injection detection to keep all of the input and output to these agents safe.

Baxter: Our Testing Center allows customers to be able to evaluate and look at their own data, see what kind of responses they’re getting from their agents, and test those agents to see if they are acting in a consistent and fair manner. We also have customizable controls that allow customers to set up their own policies to decide what data an agent can use when making a decision or not so they’re not using data that might introduce bias.

Baxter: I think one of the biggest gaps is when customers conduct testing prelaunch, but then they do not continue to monitor or they don’t continue to conduct testing. As we all know with probabilistic systems, the answer that you get today might not be the answer that you get tomorrow. Continuous monitoring of what is the performance of your agent is critical to ensuring that your agent is safe and giving your customers the best experience possible.

Go deeper:

  • Why Salesforce post-trained its own reasoning model
  • Multiplayer AI: The shift from smarter individuals to a smarter company
  • Five takeaways from Dreamforce 2026

What this article says

Something is unclear? Ask about the article — I will explain in plain words.

Do not want to dig deeper? We will sort it out for you.