40 million. That's how many AI agents Satya Nadella says are now registered on Microsoft's Agent 365 platform, and the number is only two months old. To put that in perspective, it took years for cloud computing to reach that kind of enterprise footprint, and here we have autonomous software agents multiplying inside thousands of companies before most CIOs have finished writing their governance policies. I find that gap unsettling, not because the technology is bad, but because the infrastructure to manage it is visibly playing catch-up.
You can see that catch-up happening in real time across this week's news. Google just made its agent and model evaluation tools generally available inside Gemini Enterprise, which sounds mundane until you realize what it's actually admitting: businesses have been deploying agents without consistent ways to measure whether those agents are any good. Groundcover, meanwhile, is arguing that telemetry from AI agents should never leave a company's own cloud — a position that only makes sense if you assume agents are already handling sensitive enough workflows that their behavioral logs are a security liability. And Anthropic's own research into Opus 5 suggests something more fundamental: smarter models don't just slot into existing agentic workflows, they may require entirely different context engineering. So the tooling problem isn't just "we need better dashboards." It's that the ground keeps shifting under the workflows themselves.
Ethan Mollick's new guide on delegating tasks to AI agents is a useful signal here too — not because it's groundbreaking, but because of what it implies about timing. MIT researchers are now writing practical management advice for handing off work to autonomous systems, the same way you'd onboard a new employee. That's a tacit admission that agents have crossed from "interesting demo" to "thing you need a playbook for." By the way, H2O.ai claiming the #2 spot on the FutureX leaderboard, right behind whatever's in first, tells a similar story: this is now a competitive market with real rankings, not a handful of research previews.
Here's what nags at me, though. While enterprises race to deploy 40 million agents and vendors race to build evaluation and observability layers around them, a separate report finds that Anthropic, OpenAI, and Google DeepMind — the very labs building the most capable models — score poorly on preparing for catastrophic, worst-case AI scenarios. Researchers closest to the technology are on record saying they may lose the ability to control what they're building. I don't think this is fearmongering; it's a mismatch of priorities that should concern anyone paying attention. We're optimizing agent telemetry and leaderboard rankings at a pace that outstrips our thinking about what happens if these systems fail badly, not just annoyingly. The commercial momentum is real and, frankly, impressive. But momentum without proportional safety investment is exactly the kind of gap that looks fine right up until it isn't.