Here is something that should bother anyone building with AI agents: every component in a multi-agent system can be individually correct and the system as a whole can still be wrong. Each agent's output becomes the next agent's input, so small, defensible errors compound into a bad outcome that no single step would own. I think this is the most underappreciated problem in agentic AI right now, and the argument for a deterministic control layer sitting above the agents is stronger than most vendors admit. Probabilistic components need a predictable referee.
That connects to a second, quieter problem. One study found nearly 1,000 third-party products with embedded AI running outside single sign-on, which means the identity systems that normally give security teams visibility simply don't see them. Most security tooling is built for AI an organization chose to adopt. It misses the agents that arrived inside a product someone already bought. By the way, this is the shadow IT story all over again, except the shadow tenant now makes decisions and calls tools on your behalf. If you can't see it, you can't govern it, and you certainly can't put it under a control layer.
The practical answer is evaluation, and I'm glad to see it getting serious attention. The frameworks being described test outcomes, tool use, error recovery and cost, not just whether a final answer looks plausible. That is what turns an impressive demo into a defensible release decision. Paired with sensible patterns for tool permissions, retrieval access control and human review, it starts to look like actual engineering rather than prompt-tweaking. If you're shipping agents without evals, you aren't shipping a product, you're running an experiment on your users.
Then there is the trust question. At this year's OpenAI DevDay, Sam Altman unveiled Dots, the company's new agent, and said OpenAI wants to set a new standard for privacy. Maybe so. But an agent that acts for you necessarily sees a great deal about you, and a promise made on stage is not an architecture. I'll believe it when I can inspect what the agent retains and who can reach it.
Meanwhile, the safety picture is uncomfortable from both directions. Media reports of a study say 60% of AI models failed terrorism-related safety tests once their safeguards were removed, which tells us less about the models than about how thin the guardrails on open weights really are. And Microsoft's Satya Nadella is now calling for an emergency brake on advanced AI, alongside stronger protection against model failures. It's a striking message from the CEO of a company selling this technology as fast as it can, even as Microsoft's own on-device Copilot model reportedly needs 53GB of memory, with a tested peak of 75.5GB, which is a sobering reminder of how far local AI still is from ordinary hardware.
So here is the question I'd leave you with: if the people building these systems are asking for a brake, who exactly is meant to hold the handle?