tldr; As the real-world challenges of running AI agents in production become clear, three themes emerged at We Are Developers World Congress North America: Agent containment, observability and evaluations, and agent experience.
Last month, we attended We Are Developers World Congress North America in San Jose. It was a chance to take the pulse of the developer community. We hosted a great reception with The New Stack, shared sessions on AI-assisted development and responsible AI, led a hands-on workshop built around real incidents and real code, and spent two days getting to meet developers at our booth.
Three big themes surfaced again and again throughout the week.
1: Agent containment
This wasn’t a surprise given reports of agents breaking containment and hacking systems or engaging in other unexpected and unwanted behavior. The headlines tend to point to “rogue agents” and the capabilities of the latest AI models, with some justification. But many cybersecurity experts point to failures of containment and monitoring as central culprits behind the agentic crimewave.
So it’s little wonder that agent sandboxing was a hot topic. Expanding the utility of agents depends on giving them increased access to tools and capabilities, such as giving them computing environments of their own that they can use to test and run code, access applications, and perform helpful actions for users. Figuring out how to do that safely is one of the highest stakes challenges in AI today. But containment and sandboxes are only one part of the story. You also need to keep tabs on what agents are doing, when they’re hitting guardrails, and if they’ve managed to escape containment. That’s where the next big theme comes in.
2: Observability and evaluations
A year ago, observability wasn’t top of mind for developers. I think it should have been, but I understand why it wasn’t. Agents were barely in production and everyone was still figuring why, how, and even whether to use them. Now it’s becoming increasingly clear that you can’t define predictable agentic behavior during development, so the conversation has shifted towards observing, evaluating, and adjusting agents in production.
As Laurie Voss, Head of Developer Relations at Arize, said in his mainstage talk, AI now produces far too much code for human code review. Multiple presenters covered variations of this problem. “If more stuff is getting into production without a human reading it, that means checks for “is it correct” also have to migrate from CI/CD tests to production, and that means evals,” Laurie wrote in his recap post. AI evaluations is of course exactly what Arize, the latest member of the Dynatrace family does.
There are many pieces of the AI observability puzzle, including mapping infrastructure topologies and correlating different types of data. Across several conference sessions, presenters emphasized production monitoring, evaluations, and timely runtime context. As Bob Wambach, Dynatrace VP of Market and Customer Insights, put it in his talk: “Safety is a runtime property.” Understanding what a model does in a live environment is a key part of that work.
3: Agent experience
Developer tooling used to be built for humans. That’s changing as agents become major users of databases, MCP servers, and APIs. In his conference recap, Laurie noted that Mintlify’s Han Wang said automated-agent requests accounted for 67% of traffic to the company’s hosted documentation during the period Mintlify measured.
Accordingly, we’re seeing a wave of new interfaces designed for agents. This fits into the first two themes. Agents need carefully scoped access to useful tools and data, including runtime telemetry and evaluation results that can support faster feedback. Earlier this year, we introduced BlueBox, an AI agent designed to help other agents access observability and runtime-intelligence context.
Laurie recapped a few best practices: serve documentation as clean markdown, expose your interfaces over MCP, include an llms.txt file, and add entry points built specifically for agents.
What’s still missing
One thing that didn’t get much attention: data storage. A few years ago, the rise of vector databases led to much discussion of storage for the AI era. That conversation has settled down as the tooling matures. But I suspect storage will take the stage again as teams continue to struggle with the massive amounts of telemetry data generated by agents and with querying and using heterogeneous data. It’s a core architectural concern for AI observability.
At Dynatrace, we tackle the problem with Grail, our unified data lakehouse.
The days of debating whether to use AI coding agents already feel distant. The questions now are how to contain agents, how to observe them in production, and how to design infrastructure for them. The developers getting that right will define the next decade of software.
Dynatrace, Grail, and BlueBox are trademarks of the Dynatrace, Inc. group of companies. All other trademarks are the property of their respective owners.
Dynatrace, Grail, and BlueBox are trademarks of the Dynatrace, Inc. group of companies. All other trademarks are the property of their respective owners.









