Azure rolls out AMD Helios 72-GPU racks and 500-core Venice CPUs for AI inference; H2 2026 rollout
Microsoft announced July 20, 2026 that it will deploy AMD's Helios rack-scale AI system and three new Azure VM families beginning H2 2026, targeting AI inference and data-heavy workloads. At the core: Helios packs 72 AMD Instinct MI455X GPUs (CDNA 5, 432GB HBM4 each), Venice EPYC CPUs, Pensando networking, and liquid cooling per rack. Each MI455X delivers 19.6TB/s memory bandwidth; the full rack aggregates 31TB HBM4, 260TB/s scale-up bandwidth, and 43TB/s scale-out networking. Alongside Helios, Microsoft introduces HDv2 VMs (nearly 500 EPYC cores, 4TB RAM, 32TB NVMe) for preprocessing/orchestration, HXv2 for EDA/scientific computing, and ND MI455X v7 (Helios-based inference), all purpose-built for inference not training.
The hardware pivot signals a strategic shift: while Blackwell delays extend Hopper production, Microsoft is diversifying its accelerator base away from NVIDIA mono-dependency. Helios is positioned for frontier-model serving, multi-agent systems, and RAG—workloads where inference throughput and latency matter more than training density. The HDv2's 500 cores address a recognized architectural pain point: CPU preprocessing and orchestration bottleneck GPU clusters. Microsoft is retiring older HBv2 VMs, consolidating around this new stack. ROCm software maturity (PyTorch, TensorFlow, JAX, vLLM support) is critical; compatibility validation will be necessary during GA period.
For infrastructure architects, Azure's AMD-heavy announcement reflects broader market reality: NVIDIA GPU capacity is constrained, Blackwell is delayed, and open-standard inference (ROCm, ONNX, Pensando switching) reduces cloud vendor lock-in. Helios differentiates on bandwidth-per-GPU and integrated cooling; pricing and regional availability remain TBD, but the per-token cost model suggests Azure is optimizing for sustained inference revenue rather than training capex attachment. Watch for ROCm maturity signals and customer adoption metrics in MI455X early-access programs; a successful ramp would substantially reduce Azure's NVIDIA exposure heading into 2027.
Sources
- Primary source
- Windows News: Azure AMD Helios
“Microsoft is bringing AMD's Helios racks with 72 GPUs and new VMs with nearly 500 cores to Azure in 2026. Each MI455X GPU carries up to 432GB of HBM4 memory and delivers 19.6TB/s of memory bandwidth. The full rack offers 31TB of total HBM4 and boasts up to 260TB/s of aggregate scale-up bandwidth.”