Cerebras brought a cabinet to its Supernova 2026 conference that looks more like an industrial cooling unit than a computer from the outside. But inside are three dinner-plate-sized silicon wafers, and the company claims it can generate text up to thirty times faster than systems built on graphics cards.
What the CS-4 Is
Cerebras unveiled the CS-4, the fourth generation of its system and the first to integrate three wafers into a single cabinet instead of one. Each wafer uses a WSE-3 Turbo processor with four trillion transistors, 900,000 cores, and 44 GB of on-wafer memory. The entire cabinet delivers 750 PFLOPS of performance with throughput of 7.2 Tb/s, while inter-wafer latency has fallen from the original five milliseconds to two microseconds.
What is key, however, is what you will not find inside: the silicon itself is not new. The WSE-3 Turbo is merely a tuned version of the chip introduced in 2024, with the same number of transistors, cores, and memory. The company has roughly doubled the clock frequency thanks to increased power delivery and more efficient wafer cooling.
Unrivaled Speed
When testing the GPT-OSS-120B model, the cabinet generated more than 4,400 tokens per second for a single user, roughly double the previous generation. In a clear comparison, Cerebras showed that what the CS-4 system can accomplish in a single second takes a graphics-card-based system a full thirty seconds.
Analyst firm SemiAnalysis estimates speeds of around four thousand tokens per second for the largest models, while Nvidia's Blackwell chips achieve only one hundred to two hundred. This is not a percentage difference, but an entirely new category of user experience—and one that is not merely theoretical. OpenAI recently launched the Ultrafast plan, where Cerebras infrastructure powers the GPT-5.6 Sol model at up to 750 output tokens per second, fourteen times faster than standard processing. In artificial intelligence, speed means productivity, said the company's CEO and co-founder Andrew Feldman.
Changes to Power Delivery
The most significant element of the entire system is power delivery. Cerebras moved the voltage converters from fifty millimeters away to just half a millimeter from the processor, virtually eliminating motherboard-level losses and allowing it to deliver twice as much power to the wafer. The identical transistors therefore operate faster simply because the cabinet provides them with better power delivery and cooling.
The company developed a module called the Wafer-Scale Backpack around this solution. Power delivery, direct liquid cooling, high-speed input and output, and control electronics form a single compact unit with half as many components as the previous version and a significantly higher proportion of automated manufacturing. According to Cerebras, deploying the system therefore takes just hours instead of the days previously required.
The cabinet's power consumption ranges from 120 to 140 kW. That is roughly half the figure expected from the upcoming AMD Helios and Nvidia Vera Rubin systems, even with three wafers inside the system.
The Interconnect That Determines Everything
The new input/output module has doubled throughput to 2.4 Tb/s per wafer, uses the standard RoCE v2 Ethernet protocol, and allows wafers to be connected both within and between cabinets without requiring a switch. For larger deployments, Cerebras has implemented Arista Etherlink switches, enabling the system to scale to models with more than fifty trillion parameters and process workloads from third-party hardware. The prompt prefill phase can be handled by AMD Helios cards or AWS Trainium chips, while Cerebras then handles the actual text generation.
Weaknesses
The stated thirtyfold difference comes from a comparison selected by Cerebras itself. The record result of 4,400 tokens also lacks configuration details such as the number of wafers, batch size, context length, and number of concurrent users. Moreover, this record, measured on a model with 120 billion parameters, does not provide enough information about the operating economics of trillion-parameter models, which are the reason the entire system was developed in the first place.
On-wafer memory capacity remains at 44 GB, so the largest models must be distributed across a large number of wafers and cabinets, placing enormous demands on interconnect quality. Analysts at Futurum Group also point out that doubling power density poses a significant manufacturing risk because twice as many watts flow into the same silicon, and the wafer cannot be replaced in the event of a failure. Manufacturing is handled by three contract partners simultaneously. SemiAnalysis also estimates that performance-per-watt efficiency has improved only slightly compared with the previous generation, which does not align with the manufacturer's claim of a tenfold increase.
Verdict
From an engineering perspective, the CS-4 system is an extraordinary achievement, albeit in a different area than the company itself emphasizes most. It is not a new processor, but an innovative cabinet capable of extracting twice the performance from an existing chip. Users who require answers faster than a person can read currently have no better option. However, those primarily concerned with the cost of generating a single token on large-scale models should wait for independent benchmarks with disclosed configurations.
How the Company Is Performing in the Market
Cerebras reported second-quarter revenue exceeding $209 million, twice as much as in the same period last year. Cloud services revenue rose to $126 million, while hardware sales fell to $54 million. The company valued its unbilled contracts at $25.4 billion, and management expects revenue to at least triple in 2027. By the end of 2027, it plans to have more than 600 MW of contracted data-center capacity and will provide European customers with access through the London-based technology company Callosum. The share price fell by 12.7 percent after the CS-4 system was unveiled, wiping out Monday's fifteen-percent gain.
Sources: finance.yahoo.com and wccftech.com



