Technology

Cerebras Claims 30x GPU Speed With Three-Wafer CS-4

Somewhere in Sunnyvale this week, engineers finished packaging the largest chips anyone makes — entire silicon wafers, each one the size of a dinner plate — into a machine built three to a rack. Cerebras introduced the CS-4 on Monday, and on Friday night chief executive Andrew Feldman gave the internet its guided tour. [1]

The claim attached to the machine is audacious enough to demand arithmetic. In a head-to-head run of the open model GPT-OSS-120B, given identical prompts, the CS-4 sustained more than 4,400 tokens per second for a single user; against GPU systems running the same work, Cerebras puts that at up to thirty times faster. That is the whole basis of the headline number — one model, identical inputs, tokens per second per user — and it arrives with the honest footnote that actual throughput varies by architecture, context length and precision. [1] The multiplier survives only as long as the workload does, which is precisely the caveat X spent Friday night litigating.

The rest of the specification rewards attention anyway, because it describes a different philosophy of computing. Where Nvidia stitches thousands of small GPUs into clusters that behave like one machine — if the network cooperates — Cerebras prints vast circuits onto intact wafers. Three of the new Wafer Scale Engine Turbo chips give a rack 750 petaflops of compute, memory bandwidth of 129.6 petabytes per second, and wafer-to-wafer latency down to two microseconds. There is no interconnect story because there is almost nothing to interconnect. The company says up to twice the CS-3's speed, ten times its throughput per watt, capacity for models beyond fifty trillion parameters, and first shipments before the quarter ends. [1]

The timing is the strategic fact. Five days from now Nvidia reports earnings against a consensus near $91 billion in quarterly revenue, the single number on which a meaningful share of the AI buildout's borrowed money implicitly rests — $220 billion of hyperscaler debt has already met its first indigestion this month. [2][3] Into that nervous window walks a public challenger whose pitch is not more chips but different economics per token: fast tokens earn more, and this machine claims the fastest tokens available.

Cerebras can afford the swagger because it has been paid in advance. The company went public last spring at a valuation near $56 billion and briefly touched double that on day one; its backlog runs past $24 billion, anchored by a commitment worth more than $20 billion to supply OpenAI with inference capacity over three years. [3] Feldman, who has spent a decade arguing that inference — not training — is where computing value migrates, now sells the thesis with hardware attached.

Skepticism remains reasonable and abundant. A vendor benchmark is a marketing document until customers reproduce it; SemiAnalysis founder Dylan Patel vouched for deployability rather than endorsing every number; and wafer-scale computing has buried optimistic predecessors before. [1] But notice what would have been unthinkable two years ago: a listed company shipping wafer-scale systems at all, with a frontier lab contract behind it, arriving the same season investors started asking hard questions about who pays for AI's infrastructure.

Whatever the CS-4 ultimately proves, the question it forces is the healthy kind. If speed per user really becomes the product, the industry's center of gravity shifts toward whoever owns latency. The wafers will answer before the arguments do. [1]

-- KENJI NAKAMURA, Tokyo

Get the New Grok Times in your inbox

A weekly digest of the stories shaping the timeline — delivered every edition.

No spam. Unsubscribe anytime.