Nvidia’s real moat isn’t the GPU anymore, and the company all but said so on its earnings call this week, just not in so many words. Jason Hardy, Nvidia’s VP of storage technology, spent part of last week telling reporters about a chip that isn’t the GPU at all. Vera, the CPU half of the new Vera Rubin rack, is the one Hardy credits with getting Nvidia’s storage and memory operations moving roughly three times faster than they otherwise would. Almost nobody outside the company had heard much about Vera before this week. That’s the tell.
I wrote back in June that Nvidia’s moat was CUDA, twenty-two years of developer tooling nobody wants to rewrite for AMD’s ROCm stack. In July, AMD’s own Andrew Dieckman told a room at Advancing AI that CUDA is a non-event now, that customers write against vLLM and SGLang and the layer that used to lock people in sits three rungs below where the actual work happens. I gave him more credit for that than I expected to going in. AMD backed the claim with something real too: a rack, seventy-two MI455X GPUs, and a partnership with Cerebras to split prefill and decode across two completely different architectures, which is AMD publicly conceding its own flagship chip is the wrong shape for half of inference.
None of that made Nvidia’s week read like a company under siege. It read like a company that had already relocated. Vera Rubin isn’t a GPU with a fancy name, it’s a full rack: the Rubin GPU, the Vera CPU, a Groq LPX inference accelerator (the same Groq Nvidia paid twenty billion dollars for last December specifically because it was the most credible standalone inference chip on the market), plus dedicated storage and networking units, all wired together over NVLink-C2C. Hardy’s pitch is that none of that compute matters if the data can’t reach the GPU fast enough, and that Vera’s whole job is making sure it does. Reading the transcript, I kept having the same reaction TechCrunch’s reporter had after talking to Nvidia’s people directly: this isn’t a GPU company patching around a memory bottleneck. It’s a systems company that happens to also sell you the GPU.
I keep landing on the same shape every time I look closely at this industry. CoWoS was supposed to be TSMC’s leverage over Nvidia, until it turned out Nvidia had locked up more than half of TSMC’s 2026 CoWoS capacity and was perfectly happy watching TSMC hand the rest to OSAT partners. Nvidia arranging five hundred billion dollars in financing with Apollo, Blackstone and Goldman was never really about helping customers afford GPUs so much as making sure the capital kept flowing toward Nvidia’s own order book no matter who signed the check. Every time a layer of this stack threatens to turn into a commodity, cheap chips, cheap packaging, cheap credit, Nvidia has already built the next layer up and made it proprietary. Orchestration is just the newest floor.
What actually gives me pause here isn’t the GPU competition anymore, and I don’t think it should worry Nvidia much either. OpenAI built its own Jalapeno chip around the identical insight: minimize data movement, keep the whole workload inside one connected system. The company said as much in its own writeup, and got a real benchmark win over a shipping Blackwell system out of the approach. If a rival frontier lab is independently landing on “orchestration is the bottleneck, not FLOPs,” that tells me the constraint is real and not just a talking point Nvidia handed a reporter. The difference is that Nvidia is selling the orchestration hardware to every other company building a data center, while OpenAI built one specifically shaped chip for its own workload. That’s not a small difference. It’s the gap between owning a layer of the market and owning a single deployment.
I don’t expect this to be a permanent win any more than CUDA was. Hyperscalers building their own chips are going to start building their own orchestration silicon too, because the same economics that pushed Amazon and Google toward Trainium and Tensor apply just as hard one layer up. But right now, in the specific window where gigawatt-scale data centers are getting built faster than anyone can actually run them efficiently, Nvidia is the only vendor selling a rack engineered end to end for exactly that problem. AMD proved this year it can match Nvidia chip for chip. Matching Nvidia rack for rack, storage stack and networking fabric and inference accelerator included, is a much longer project, and AMD hasn’t announced anything that touches it yet.
Sources
- TechCrunch, Nvidia’s AI advantage is moving beyond the GPU, August 29, 2026
- OpenAI, Jalapeno: first results, August 2026
- PC Gamer, AMD calls CUDA a ‘non-event’, July 23, 2026
- TrendForce, TSMC Reportedly Expands Outsourcing of Key CoWoS Front-End Step to OSATs, August 5, 2026
- CNBC, Nvidia teams up with Wall Street asset managers on $500 billion AI infrastructure push, August 10, 2026