The Thai Ministry of Finance espionage story deserves more attention than it's getting, because it marks a shift I've been waiting for and dreading in equal measure. Attackers reportedly deployed Hermes, an autonomous open-source AI tool, in what's being described as an unrestricted "YOLO mode" to conduct the operation. Not a chatbot helping draft phishing emails — an agent given the run of the place, making its own decisions about how to breach a government ministry. We've spent two years talking about agentic AI as a productivity story. This is the reminder that autonomy cuts both ways, and the same properties that make agents useful — the ability to chain actions, adapt, and operate without constant human sign-off — make them dangerous in the wrong hands.
What's striking is how the defensive and offensive stories are converging on the same week. Wiz's Atlas, an autonomous agent built for vulnerability research, just topped the CyberGym rankings by verifying every finding with a working, real-world exploit rather than a theoretical writeup. That's genuinely useful — security researchers get an agent that doesn't just flag a hunch but proves it. But it's also a mirror image of what apparently happened in Thailand: an AI system operating with minimal human oversight, finding weaknesses, and acting on them. The tooling for defense and offense is converging, and the gap between "red team agent" and "espionage agent" is mostly a matter of who's holding the leash — and how tightly.
This is why the supply chain conversation matters more than it sounds like it should. Experts are now recommending version pinning, package firewalls, and least-privilege access specifically because AI agents can autonomously pull in vulnerable dependencies without anyone reviewing the choice. It sounds like dry infrastructure advice, but it's really the same problem showing up one layer down: agents making consequential decisions faster than humans can review them. AINGENS is making a related point about agentic AI in life sciences — that autonomy in regulated, high-stakes workflows has to be constrained by verifiable evidence and traceability, not just good intentions. I think that's the right framing for agentic AI generally right now, not just in pharma. The question isn't whether agents should be trusted with more, it's what verification looks like at scale, because right now most organizations are nowhere close to answering that.
Meanwhile the corporate world keeps racing in the opposite direction, embracing autonomy faster than governance can keep up. Meta is wiring its AI directly into employee inboxes and calendars, letting it act rather than merely suggest. Microsoft Japan just launched Cowork and Scout, autonomous agents that INPEX expects will save it over ¥2 billion a year. HubSpot is building tools just to manage the sprawl of agents companies are already deploying. All sensible business moves individually. Collectively, they're expanding the attack surface at exactly the moment we've seen what a malicious actor can do with the same category of tool. By the way, nobody seems to be asking who's auditing all these new agents once they're live — that's probably the story of 2027, not 2026.