Sean Lie on why yesterday's fast is the new batch mode
Cerebras co-founder and CTO Sean Lie argues that wafer-scale speed is redefining inference: yesterday's fast is becoming batch mode, speed compounds into smarter agents, and the real frontier is moving to integration beyond the single chip.
Yesterday’s “Fast” Is the New Batch Mode
<strong>The definition of “fast” keeps sliding:</strong> token rates once considered fast now feel like overnight batch jobs, while real-time becomes the new bar.
what used to be batch and offline applications start to become real time
Speed Compounds Into Intelligence
<strong>Faster tokens are not just a smoother interface:</strong> they buy more agentic loops and more reasoning inside the same time budget, which produces measurably more capable agents.
this flywheel of not just making models you know faster but using that speed to make them more intelligent
From “You Guys Are Insane” to Sold Out
<strong>Wafer-scale moved from a technology demo to sold-out capacity:</strong> the first chip only had to prove the idea was possible, and the job now is scaling it.
this is the future and most of them were just like you guys are insane this is not possible
The 30-Billion-Parameter Tell
<strong>Memory capacity, not raw speed, is the real constraint:</strong> a small-memory SRAM chip can only benchmark a small model because it lacks the memory to hold frontier weights at all.
one of our chips has you know order 100 times more memory than one of their chips right so you have two orders of magnitude difference in scale kind of for free
Treat the Data Center as One Chip
<strong>Inference is no longer one workload but many:</strong> prefill, decode, KV-cache loading, attention, and expert routing can each get purpose-built hardware inside one disaggregated system.
it's as if you're like thinking about the entire data center as if it's like one computer. It's like one chip that you're trying to figure out
We’re All Running Models Designed for Nvidia GPUs
<strong>Today’s models are shaped for one specific Nvidia GPU:</strong> even slight architectural adjustments for the target hardware can unlock large additional speedups.
if you start to then open up the possibility of adjusting that model architecture even slightly, you can get massive gains
The Frontier Is Outside the Chip
<strong>No single chip can hold a frontier model,</strong> so the next big gains come from integration across memory, interconnect, and packaging at the package, node, and rack.
the applications, the problems, the models are now so large that you can't just do anything on a single chip. So everything comes down to the integration
The Real Wafer-Scale Boss Was Power
<strong>The hard part was not wiring the wafer but feeding and cooling it:</strong> Cerebras expects 3D DRAM stacking to run into the very same wall.
what turned out to be, you know, the biggest enablers was ultimately how do you power it? How do you cool it