aiexpert
Home / News / Brief
Chips · Aug 10, 2026, 07:32 AM · 6 sources

AMD Instinct MI455X ships with 432GB HBM4, 40 PFLOPs FP4, rivals NVIDIA Rubin in Helios rack-scale platform

AMD officially launched the Instinct MI455X, its CDNA 5 data-center GPU designed for frontier AI training and inference, at its Advancing AI 2026 event in late July. The accelerator packs 432GB of HBM4 memory (50% more than NVIDIA's Rubin at 288GB) with 23.3 TB/s memory bandwidth, 40 PFLOPS of FP4 compute, and 320 billion transistors across a 2nm/3nm chiplet design. The MI455X powers the Helios rack-scale AI system, which aggregates 72 GPUs into a unified domain delivering 2.9 exaFLOPS at FP4 and 31 TB of shared HBM4 memory at 1.67 PB/s aggregate bandwidth. AMD explicitly positioned the accelerator as a direct competitor to NVIDIA's GB300 NVL72 and next-gen Vera Rubin platforms.

The MI455X uses TSMC's advanced CoWoS-L packaging with four Accelerator Complex Dies (8 XCDs total) on 2nm bonded to two Fabric/Cache Dies on 3nm, each connected to six HBM4 stacks via a 2048-bit memory interface. CDNA 5 makes significant architectural changes: it drops the Wave64 compute model for native Wave32 (lower latency for agents), adds a 96MB local L2 cache per die, introduces a Tensor Data Mover for GPU-memory transfers, and supports up to eight spatial partitions via NUMA. The Helios rack uses open standards—OCP, UALink over Ethernet, and UEC—rather than proprietary interconnects, and connects via 260TB/s scale-up and 43TB/s scale-out bandwidth.

For architects: MI455X's 50% memory capacity advantage over Rubin matters most for long-context inference and large-batch KV caching where staying resident in GPU memory saves communication overhead. Benchmark gaps remain: AMD claims 15% higher throughput than Rubin on internal tests, but real-world performance depends on ROCm optimization maturity and quantization strategies that haven't been independently validated yet. Helios volume deployments are expected in H2 2026. The ecosystem play is significant—AMD is committing to yearly accelerator releases (MI500 in 2027, MI600 in 2028), signaling faster competitive cadence than NVIDIA's historical pace. For teams building cloud-native inference infrastructure, the open-standards bet on UALink and OCP tooling may reduce Nvidia lock-in risk.

Sources

Everything this brief rests on
  1. 01 Primary source hothardware.com
  2. 02 AMD Ships Helios AI Rack With EPYC 9006 CPUs, Instinct MI400X GPUs hothardware.com “The MI455X moves to the CDNA 5 architecture which seems to be quite a radical departure from the previous generation.”
  3. 03 AMD Instinct MI455X GPU features 432GB HBM4 and 23.3 TB/s memory bandwidth videocardz.com “The accelerator includes 432GB of HBM4 memory and offers up to 23.3 TB/s of peak memory bandwidth. AMD lists 320 billion transistors and a 2nm process node.”
  4. 04 xenospectrum.com xenospectrum.com “Helios repeatedly arranges compute trays, each loaded with four MI455X units, to fit 72 GPUs into a single rack. The host uses the 6th-generation EPYC Venice.”
  5. 05 tomshardware.com tomshardware.com “All told, the MI455X is AMD's most compelling Instinct product yet. It delivers a massive generational performance improvement.”
  6. 06 chipsandcheese.com chipsandcheese.com “It is AMD's first GPU designed for rack-scale AI deployments and is based on the new CDNA5 architecture with major changes to the compute unit and SoC.”