AUTO-UPDATED

Nvidia paid $20 billion for SRAM decode - AMD just partnered for it instead

AMD and Cerebras Systems have announced a technical partnership to create a disaggregated inference solution, combining Helios rackscale systems and Wafer-Scale Engines for improved energy efficiency by 2026.

Key Points

  • The partnership pairs AMD Helios racks for prompt processing with Cerebras Wafer-Scale Engines for token generation.
  • Internal testing indicates the combined configuration delivers five times higher tokens per second per watt compared to standalone WSE setups.
  • The solution is scheduled to become available to customers through the Cerebras Cloud in the second half of 2026.
  • Testing utilized the Kimi 2.6 1T model, a mixture-of-experts design that operates natively in INT4.
  • The collaboration allows AMD to enhance its inference capabilities without the multibillion-dollar licensing costs associated with alternative SRAM decode technologies.

Why it Matters

This partnership represents a strategic effort to optimize AI inference efficiency by leveraging the specific hardware strengths of two different architectures. By offloading resource-intensive prompt processing to AMD, the companies aim to improve performance sustainability and provide a competitive alternative to existing market-leading infrastructure.
TechRadar Published by Rahimnoorali11@gmail.com (Rahim Amir) , Rahim Amir
Read original