AMD, Cerebras team on disaggregated AI inference; claim 5x efficiency gains
AMD and Cerebras announced a technical partnership to deliver a disaggregated AI inference platform combining AMD Helios rackscale systems with Cerebras Wafer-Scale Engine (WSE). Unveiled at Advancing AI 2026 on July 23, the joint solution pairs AMD's high-throughput Helios rack (72 MI455X GPUs per rack, 31TB HBM4) with Cerebras' ultra-low-latency WSE-3 (900,000 cores, 44GB on-die SRAM). Together, they claim up to 5x higher tokens per second per watt versus Cerebras WSE-only configurations.
AMD Helios handles the context/prefill stage—processing prompts and large context windows at high throughput—while Cerebras WSE accelerates the memory-bandwidth-intensive decode/token-generation stage with sub-10ms latency per token. This disaggregated approach mirrors NVIDIA's Rubin CPX strategy but inverts the specialization: AMD targets throughput, Cerebras targets latency. The result is positioned for applications prioritizing real-time interaction: copilots, live agents, autonomous workflows, and scientific discovery.
Cerebras plans to deploy AMD Helios in its own data centers, with the joint solution expected to be available initially through Cerebras Cloud in H2 2026. For architects: this partnership addresses a real bottleneck—NVIDIA's Rubin also pursued disaggregation for the same reason. The question for buyers is feature parity and software ecosystem. AMD and Cerebras must demonstrate that mixed inference across Helios + WSE matches the ease of single-vendor stacks, particularly for teams migrating from NVIDIA.