Cerebras Unveils CS-4
Key point
Cerebras launched the CS-4, featuring three wafers and delivering up to 30x faster inference performance compared to GPUs.
Details
Cerebras Systems unveiled its fourth-generation wafer-scale system, CS-4. The CS-4 is a system housing three units of the new WSE-3 Turbo processor in a single rack, achieving inference speeds up to 30x faster than GPUs and up to 2x faster than the previous generation CS-3, measured in tokens per second per user (TPS/user). Total token throughput increased by up to 10x compared to CS-3 within the same power budget.
The WSE-3 Turbo maintains the same transistor count and SRAM capacity as the existing WSE-3 but doubles compute performance and bandwidth by modifying the rack design to shorten power delivery distances and increase operating frequency. The CS-4's system memory bandwidth reaches 129.6PB per second, thanks to an architecture that reads weights from on-wafer SRAM.
The new Nexus rack-scale platform separates compute, power, and I/O into independent modules to support flexible scaling. Notably, it is designed to maintain interactivity in large models with 10 trillion parameters by reducing I/O latency to 2 microseconds. Cerebras emphasized that this system can significantly increase the number of inference and verification steps in agent systems, thereby enhancing AI productivity.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.