AMD is partnering with Cerebras, a high-speed AI inference system provider, to ship ultra-low latency AI systems in the second half of 2026.



Cerebras, a company developing systems to accelerate AI inference, has partnered with AMD. Cerebras plans to deploy AMD's rack-scale solution,

Helios , in its data centers to enhance its AI inference capabilities.

AMD and Cerebras Announce Industry-Leading Ultra-Low-Latency and High Throughput AI Inference Solution - AMD Newsroom
https://newsroom.amd.com/news/aai-2026-cerebras-inference/




On July 23, 2026, AMD and Cerebras announced a partnership 'to advance workload optimization approaches for ultra-low latency AI inference infrastructure.'

Cerebras plans to implement AMD Helios and combine it with its own AI chip, ' Wafer-Scale Engine ,' to offer a new 'separated AI inference solution.'

According to AMD, while maximizing token generation is crucial in AI inference processing that handles large amounts of information, coding and agent-based workflows require fast response times, meaning the optimal behavior for the user varies depending on the situation.

A new segregated AI inference solution divides these processes into two parts and optimizes each independently. Helios achieves ultra-high throughput and handles prompts and large context windows, while the Wafer-Scale Engine takes over by accelerating token generation, which is memory bandwidth-intensive, with ultra-low latency. This is expected to achieve the ultra-low latency required for cutting-edge AI applications while dramatically improving throughput and efficiency.



Lisa Su, Chairman and CEO of AMD, said, 'AI inference is becoming one of the largest infrastructures in the AI field, and as it diversifies, a more flexible approach is required. AMD Helios delivers industry-leading performance and scale for a wide range of inference workloads. Our collaboration with Cerebras extends that leadership to the most latency-sensitive applications and builds a powerful new platform for real-time agent-based AI.'

Andrew Feldman, CEO and co-founder of Cerebras, commented, 'The demand for ultrafast inference is expanding at an unprecedented pace. Cerebras offers the world's fastest ultra-low latency inference. Our partnership with AMD presents a great opportunity to bring that performance to even more customers.'

in AI,   Hardware, Posted by log1p_kr