AMD and Cerebras Systems have announced a technical partnership to create a disaggregated inference solution, combining Helios rackscale systems and Wafer-Scale Engines for improved energy efficiency by 2026.
Key Points
- The partnership pairs AMD Helios racks for prompt processing with Cerebras Wafer-Scale Engines for token generation.
- Internal testing indicates the combined configuration delivers five times higher tokens per second per watt compared to standalone WSE setups.
- The solution is scheduled to become available to customers through the Cerebras Cloud in the second half of 2026.
- Testing utilized the Kimi 2.6 1T model, a mixture-of-experts design that operates natively in INT4.
- The collaboration allows AMD to enhance its inference capabilities without the multibillion-dollar licensing costs associated with alternative SRAM decode technologies.