AMD launches Helios rack; 15% more compute, 30% better tokens/dollar than Vera Rubin
AMD officially launched Helios, its first rack-scale AI system, at Advancing AI 2026 in San Francisco on July 23. The system combines 72 AMD Instinct MI455X GPUs with 6th-gen EPYC Venice CPUs and Pensando networking into a double-width rack (4OU) delivering 2.9 exaflops of FP4 inference performance. AMD claims Helios delivers 15% better compute performance and 50% greater high-bandwidth memory (HBM4) capacity than NVIDIA's Vera Rubin NVL72, while offering 30% more inference tokens per dollar.
Helios is priced between $5–5.5 million per rack and enters production immediately, with early customer deployments beginning in Q4 2026. Major customers announced include Microsoft Azure, Meta (1 GW committed to Helios), OpenAI (12 GW of AMD GPUs across OpenAI and Meta), Oracle, Anthropic (2 GW partnership with engineering collaboration to use Claude for ROCm optimization), and Tata Consultancy Services. Each rack provides 31 terabytes of HBM4 memory and 1.7 petabytes/second of aggregate bandwidth.
AMD's move directly challenges NVIDIA's >95% data center GPU market dominance; AMD currently holds ~4.5%. Helios represents AMD's most integrated play yet—bunching silicon, networking, and software with strategic partnerships on inference workloads (via Cerebras) and model optimization (via Anthropic, OpenAI). For architects planning multi-vendor AI infrastructure and evaluating total cost of ownership over time, AMD Helios signals credible rack-level competition in 2026–2027, with supply chains now diversifying beyond NVIDIA monopoly.
Sources
- Primary source
- cnbc.com
“AMD launched Helios, combining 72 MI455X GPUs with EPYC CPUs; CEO Lisa Su claimed 15% better compute than Vera Rubin, 50% more HBM capacity, 30% more tokens/dollar; Microsoft Azure adopting it alongside Meta, OpenAI, Oracle”
- qz.com
“Anthropic partnership to deploy 2 GW of MI455X GPUs; OpenAI bringing Helios online Q4 2026; Microsoft Azure deploying in second half of 2026”
- ir.amd.com
“Helios now in production; 2.9 exaflops FP4, 1.4 exaflops FP8, 31 TB HBM4, 1.7 PB/sec bandwidth; OpenAI and AMD optimizing GPT-class workloads with Triton + ROCm”