On August 18 (local time), Cerebras introduced its latest-generation AI accelerator, the CS-4, asserting that it can achieve inference speeds up to 30 times swifter than those of GPU-based solutions. The CS-4 is outfitted with three cutting-edge WSE-3 Turbo wafer-scale engines, providing an impressive AI computing capacity of up to 750 PFLOPS and a memory bandwidth of 129.6 PB/s. During the GPT-OSS-120B test, the CS-4 demonstrated its prowess by generating over 440 tokens per user per second—a speed up to 30 times greater than that of GPU solutions. Additionally, it boasts a per-watt throughput up to 10 times higher than its predecessor, the CS-3, and is capable of supporting models with over 50 trillion parameters. The CS-4 is slated for release in the third quarter of this year.
