Dan Bikel | October 1, 2026
When thinking about transparency in AI, many organizations look closely at the model they’re using and the outputs it generates. That’s a good start, but it has blind spots that could come back to bite you if your AI takes actions or produces results that damage your security, business or brand. Enterprises have moved beyond simple generative AI and are now working with agentic AI systems that touch many different data sources, connected services, and autonomous workflows.
Agents are powered by models, but they act based on much larger systems consisting of prompts, tools, retrieval, memory, permissions, guardrails and other agents. Transparency for agentic AI has to extend from the underlying language model through the full system the agent operates in.
Model transparency is necessary, but not enough
When we discuss model transparency, we mean more than just reviewing the model card released with it. Enterprises need the equivalent of a mechanic’s report: useful information about model architecture, training data, evaluations, limitations and intended behavior.
The latest Stanford Foundation Model Transparency Index offers a useful reality check.
In its December 2025 assessment, Stanford evaluated 13 major foundation-model developers against 100 transparency indicators. The average score was 41 out of 100, down 17 points from the previous edition. WRITER’s Palmyra X5 scored 72, the second-highest score in the study behind IBM’s Granite 3.3 at 95.
That distinction matters because model transparency is not the same thing as model performance. Stanford’s framework examines whether developers disclose information across areas including training data, model evaluation, downstream use, supported agent protocols and intended behavior. Stanford’s report on WRITER, for example, notes disclosures around data acquisition, the fact that customer data is not used to train Palmyra X5, supported MCP and A2A protocols and model capability evaluations.
For us, the broader principle matters more than the ranking itself. Enterprises should expect model providers to document how their models are built, tested and intended to behave. We believe that obligation becomes more important as models become more powerful, embedded inside agents capable of taking action with real-world consequences.
But model transparency is a foundation, not the entire framework. A model card tells you what the model is. It does not tell you what the agent built on that model will do, i.e., which tools it can call, what data it can access, what guardrails constrain it, or what context shapes its decisions. For that broader lens, you need a framework.
The Agentic Compact: a framework for agentic transparency
The Agentic Compact is WRITER’s framework for evaluating agentic AI systems. It extends transparency from the underlying language model through the way the agent is designed and operates, covering model selection, the system around the model, the context that shapes behavior and the evidence of the agent’s actions. The Compact spans six areas: systemic safety and containment, foundational transparency, actionable explainability, continuous observability, workforce enablement and education, and the human mandate. This piece focuses on foundational transparency: the visibility enterprises need before an agent ever goes into production.
For enterprise leaders, the Compact poses four questions:
1. Which model is this agent using, and why?
The right model depends on the task: reasoning quality, latency, cost and security requirements all narrow the field, and enterprises need enough information about a model’s capabilities and limitations to know whether it’s appropriate for the work being delegated to it.
2. What is the harness around that model enabling or constraining it to do?
The harness—the instructions, tools, permissions, and guardrails surrounding the model—determines what the agent can and cannot do, and two systems using the same model with different harnesses are not meaningfully the same agent.
3. What context is shaping its behavior?
Agents draw on internal knowledge, business rules, customer data, and persistent memory, and the provenance of that context, viz., where it came from, whether it’s current, and who can correct it — is becoming as important as the provenance of the model itself.
4. What evidence do we have about the actions and decisions the model itself is making and why?
Enterprises need far more than just a model card describing the main LLM. Rather, enterprises need action-level evidence: logs of what the agent did, which tools it called, what data it accessed, and why it made the decisions it made.
Transparency must go beyond disclosure, providing the visibility required to make informed choices and retain control from model choice through harness selection, data connections and the steps that led to the final output.
Which model is this agent using, and why?
Regardless of benchmark scores or slick marketing campaigns, there is no universal “best” model. There is only the model best suited for a particular task and set of constraints.
For one workload, reasoning quality may matter most. For another, latency may be the binding constraint. For a high-volume workflow, the decisive issue may be cost. In regulated environments, security and data-handling requirements can narrow the field further.
The Agentic Compact starts with model selection and provenance. Leaders should understand what sits underneath an agent, what is known about that model’s capabilities and limitations, and whether it is appropriate for the work being delegated to it.
When we introduced Palmyra X6, we also updated the WRITER platform so teams could access multiple models rather than being limited to the Palmyra family. Palmyra X6 remains the default, but administrators can enable models from other providers, and enterprises can also bring models through infrastructure including AWS Bedrock, Microsoft Azure and NVIDIA NIM.
While a platform that lets you choose between models is an excellent start, it is not enough. Enterprises need enough information to understand why that model is appropriate for the workload and whether another model would perform the job more efficiently.
That last consideration is of growing importance as AI spend balloons and tokenomics, like cloud costs before it, becomes its own discipline. Every C-suite leader who owns an AI budget is also, whether they asked for the job or not, becoming responsible for a token budget.
The cheapest model is not necessarily the cheapest way to complete a task. A lower-cost model that takes repeated attempts, calls more tools, or consumes far more tokens may produce a higher cost per finished task than a more capable model. Conversely, using the most expensive frontier model indiscriminately can turn routine work into an unnecessarily large bill.
My colleague Matan-Paul Shetrit’s recent piece on tokenomics makes the useful distinction here: enterprises should optimize for the economics of completed work rather than treating model intelligence or token price as isolated metrics. Multi-model support gives enterprises the ability to act on what they learn about performance, cost, security and efficiency. It also avoids vendor lock-in, ensuring the data and workflows built as part of a company’s ongoing AI journey are not siloed to a single provider.
What harness surrounds the model?
Choosing the right model is only part of what matters.
The same model can behave very differently depending on the harness around it: the instructions it receives, the tools it can call, the retrieval and memory systems feeding it information, the permissions it operates under and the guardrails constraining its actions.
That harness is increasingly where a large part of agent performance is determined.
When we released Palmyra X6, we also introduced major upgrades to the WRITER Agent harness. The model and agent were developed together, with the harness designed to improve how complex, multi-step work is planned and executed. The WRITER Engineering team published a white paper showing that the upgraded system reduced cost and improved task completion speed across the models it tested.
If two systems use the same underlying model but surround it with different prompts, tools, guardrails and orchestration logic, they are not meaningfully the same agent. Enterprise leaders therefore need visibility not only into which model is running, but also into what surrounds it. That is the difference between mere model transparency and full system transparency.
If your model is the engine of the car, the harness is the machinery that translates its raw capability into a tool suited to the task at hand. The next question you need to consider is, beyond the engine and it’s machinery, what fuels it. This requires you to examine the information that flows through it.
What context is shaping the agent’s behavior?
The model is only one source of intelligence inside an enterprise agent.
Agents increasingly draw on enterprise-internal knowledge, business rules, customer data, instructions, previous work and persistent memory. As that context becomes richer, its provenance becomes as important as the provenance of the model.
The Compact identifies three basic questions enterprises should be able to answer: What knowledge sources can the agent access? How is that information processed and secured? And how does the system decide which information belongs in the agent’s working context?
WRITER’s recent Enterprise Brain launch pushes that idea further. Enterprise Brain is a governed context layer designed to capture institutional knowledge, decision logic, brand and compliance standards and live business data, then make that context available across teams and agents.
Agent Memory is the system that creates, manages and retrieves remembered information for each Enterprise Brain-equipped agent. Far from treating memory only as a store of individual preferences, WRITER’s engineering approach emphasizes the kind of work-related information our enterprise customers need most: tasks, requirements, outputs, team conventions and organizational knowledge accumulated over time.
The addition of a system like Enterprise Brain with Agent Memory changes what transparency requires. We now need to be able to answer the following questions:
- Where did this context come from?
- Is it current?
- Is it personal, team-level or organizational knowledge?
- Who can correct or supersede it?
- Which agents are allowed to use it?
- What happens when two sources of remembered context disagree?
A powerful model working from stale, opaque or incorrect context is still an unreliable system. In agentic AI, context provenance is becoming as important as model provenance.
Conclusion
AI transparency used to be largely a question about models and their basic design. Today, the scope of what an enterprise needs to observe and understand is both broader and deeper, covering everything that surrounds the main LLM, the data that it acts on, and the other models and systems that store and retrieve data external to the main LLM.
Enterprises need to know which model is being used for a particular task and why. They need visibility into the harness of prompts, tools, permissions and guardrails around it. They need to understand what enterprise context and memory are shaping its behavior. Finally, they need meaningful information about the model itself.
The goal is agency through transparency: the ability to inspect the system, compare alternatives, understand what is informing its behavior, and make a different choice when necessary.
– Dan Bikel is the Head of AI at WRITER. He led the team building our latest foundation model, Palmyra X6, played a key role in helping to develop WRITER’s approach to agentic memory, and leads the team doing cutting-edge research to power products and services now and into the future. Previously, Dan worked as a senior tech lead at Meta in the GenAI org, leading the team that built agent memory and new methods for contextual intelligence. Prior to that, Dan was a research scientist at Google Research, investigating new methods for acquiring and organizing knowledge, discovering new entities, question answering, parsing queries, transcribing speech for YouTube videos and more.







