AMD's Helios racks beat Nvidia's Vera Rubin on paper — but paper's cheap
Curated by the Inblix editorial team
AMD just drew a line in the sand with Helios, its first true rack-scale AI compute platform, and on paper it makes Nvidia’s upcoming Vera Rubin look underpowered. We’re talking more teraflops, more memory bandwidth, and a design that spec-for-spec outmuscles the competition by nearly every metric. The timing is deliberate. AMD’s Epyc silicon and Instinct GPUs have been winning socket by socket in the cloud, but Nvidia still owns the AI narrative. Helios is AMD’s attempt to change that conversation from one about chips to one about complete systems — the kind hyperscalers actually buy.
The elephant in the room, of course, is that Nvidia’s Vera Rubin isn’t shipping yet either. This is a paper launch from AMD, and the company has a history of impressive specs that take longer than promised to materialize. But the architecture is real. AMD is betting on an open ecosystem approach with its Infinity Fabric tying CPUs and GPUs together, while Nvidia continues to push its proprietary NVLink. For cloud builders tired of writing blank checks to Jensen Huang, Helios offers a genuinely credible alternative — if it ships on time and performs as advertised.
Cerebras and Groq are fighting their own proxy war in this same battle. Cerebras, with its dinner-plate-sized wafer-scale chips, has partnered with AMD to create systems optimized for inference workloads where Nvidia’s LPU architecture currently dominates. It’s a classic “enemy of my enemy” play. Cerebras brings absurd memory bandwidth to the table; AMD brings the system integration and customer relationships. Together they’re targeting the inference market that Nvidia’s Grace-Hopper and upcoming Vera Rubin architectures were supposed to lock down.
The real question isn’t whether AMD can build a faster rack — it’s whether anyone other than the hyperscalers can afford it. At this level, we’re talking about systems that cost more than most companies’ entire data center budgets. AMD’s strategy hinges on convincing a handful of cloud giants that vendor diversity matters enough to qualify a second source. If Helios delivers even 80% of what’s promised, it changes the procurement calculus. But Nvidia’s CUDA moat is deep, and specs on a slide deck don’t compile code. I’ll believe it when I see it running Llama-4 at scale without catching fire.
💡 Key Takeaways
- AMD's Helios platform specs out faster than Nvidia's Vera Rubin across nearly every metric, but both systems remain unreleased — making this a battle of slideware for now.
- Cerebras is partnering with AMD to attack Nvidia's inference dominance, leveraging its wafer-scale chips' massive memory bandwidth as an alternative to LPU architectures.
- The real barrier isn't performance — it's Nvidia's CUDA software ecosystem, which keeps developers and enterprises locked in regardless of hardware specs.
- Hyperscalers may be the only customers with budgets large enough to deploy these rack-scale systems, meaning AMD's success depends on convincing just a handful of buyers.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.