r/HPC • u/healthjay • 1d ago
AMD Helios vs NVIDIA Vera Rubin NVL72: comparing two 72-GPU rack architectures
We have put together a side-by-side comparison of AMD Helios and NVIDIA Vera Rubin NVL72:
https://linuxclusters.com/articles/amd-helios-vs-nvidia-vera-rubin/
The shared 72-GPU rack boundary hides some different design bets. On paper, Helios has more accelerator memory and scale-out bandwidth and leans more heavily on open rack and fabric standards. Vera Rubin has more memory bandwidth per GPU and comes with NVIDIA's more integrated software and networking stack. The comparison covers hosts, memory, fabrics, networking, software, and rack standards, while identifying gaps in the available power, pricing, reliability, and application-performance data.
I would especially welcome corrections from people working with rack-scale systems. Which of these differences is most likely to matter in an actual deployment, and which missing numbers would you insist on seeing before procurement?
Disclosure: I am one of the writers at LinuxClusters.com