Everyone wants to own the chip underneath your AI agent, and today's news makes clear why that fight matters more than the models themselves.
Start with the silicon arms race. Apple just unveiled its M6 and M5 Ultra chips, pitched heavily around AI compute rather than the usual "faster everything" pitch. Meanwhile OpenAI is reportedly building its own inference chip, codenamed Jalapeño, specifically to reduce dependence on NVIDIA. That's a striking move from a company that has spent years being NVIDIA's most prominent customer. If OpenAI succeeds in designing hardware that beats NVIDIA GPUs for inference — the actual running of models, as opposed to training them — it changes the economics of the entire industry. Inference is where the real, recurring costs live once a model ships. Whoever controls that layer controls margins. NVIDIA isn't standing still either, and Groq has quietly built a case that its specialized chips, not the underlying model architecture, are what actually determine whether an AI agent feels usable in real time. This is the unglamorous part of the AI story that rarely makes headlines, but it's arguably more consequential than the next chatbot release: the bottleneck has shifted from "can we build a smart model" to "can we run it fast and cheap enough to matter."
That shift matters because agentic AI is where the money is supposedly moving next, and the early results are messy. Salesforce says its AI agent revenue is up more than 200%, which sounds spectacular until you notice the caveat — it's unclear how much of that shows up in the company's official order backlog, the number investors actually trust. That gap between "look how fast this is growing" and "but can you prove it on the books" is going to define a lot of AI earnings calls over the next year. Info-Tech Research Group's warning fits the same pattern: agentic AI projects aren't failing because the technology is immature, they're failing because companies are throwing budget at the wrong pilots. I find this more useful than most agentic AI hype, because it reframes the problem correctly — this is a capital allocation failure, not a model capability failure. Perplexity teaming up with NVIDIA to ship a local AI agent is a small but telling counterpoint: running agents on-device, rather than routing everything through the cloud, is one practical answer to both the cost problem and the latency problem Groq is betting on.
By the way, there's a quieter thread running underneath all of this that deserves more attention than it gets: the idea of "instrumental succession," the slow, undramatic handover of decision-making from humans to AI systems. No one signs a single document authorizing this. It happens one delegated task at a time — approve this agent to send the email, approve this agent to allocate the budget, approve this agent to choose the vendor. Alice raising $140 million to secure AI models against tampering and theft is a sign that at least some investors take the stakes of that handover seriously, even if the framing is mostly commercial rather than philosophical.
So which race actually decides who wins the agent economy — the chip war, the capital discipline problem, or the quiet accumulation of authority nobody explicitly granted?