Cerebras' 'CS-4' rack-scale solution delivers AI up to 30 times faster than GPUs.

Cerebras, a company developing systems to accelerate AI inference, has unveiled its rack-scale solution, ' CS-4 .' It is touted as being able to perform inference processing up to 30 times faster than a simple GPU system.
Product - System - Cerebrals
Introducing Cerebras CS-4: The Fastest AI Gets Faster
https://www.cerebras.ai/blog/introducing-cerebras-cs-4
The CS-4 features three Wafer Scale Engine 3 Turbo (WSE-Turbo) processors per system, enabling each wafer to achieve up to twice the speed of the previous generation. This, combined with a completely new power supply, cooling, and I/O system, further enhances performance per wafer.

Equipped with WSE-Turbo, the CS-4 achieves inference up to 30 times faster than a simple GPU system. By enabling low-latency communication, the CS-4 can process more than 1,000 tokens per second with a model having over 10 trillion parameters. Cerebras advertised that it 'generates tokens at a speed that GPU systems simply cannot match.'

Cerebras stated, 'Until now, AI infrastructure has been forced to choose between high bidirectionality and high total throughput. CS-4 changes this. CS-4 generates tokens up to 30 times faster than conventional GPU systems, while simultaneously improving throughput per watt by up to 10 times compared to CS-3. Because faster tokens are more valuable than slower tokens, CS-4 significantly improves data center profitability by supplying more valuable and more tokens within a limited power budget. Users will enjoy a more responsive experience, and operators will gain the ability to handle more workloads.'
Furthermore, the CS-4 reduces the number of components by 50% and shortens manufacturing and deployment time by integrating the wafer, power conversion, liquid cooling, high-speed I/O, and control electronics into a compact 3D package. In particular, by separating the computing and power supply, the manufacturing process has been simplified and deployment time has been reduced from several days to several hours.

In terms of power efficiency, improvements have been made by positioning the power supply just 0.5mm away from the processor. This is nearly 100 times the distance compared to the approximately 50mm on conventional GPU boards, and this has doubled the amount of power that can be supplied while virtually eliminating power loss.

Cerebras stated, 'CS-4 reduces deployment time from days to hours and simplifies maintenance and future upgrades in hyperscale environments. Initial shipments of CS-4 will begin this quarter (by the end of September 2026).'
Related Posts:







