Amin Vahdat on Goodput, Not FLOPS
Google aims to double its token-serving capacity every six months, with as much of the gain coming from software as from hardware, and measures itself on goodput under real-world failures rather than peak FLOPS.
Goodput, Not FLOPS
The number that matters is not a chip's theoretical FLOPS but goodput: the useful work that actually reached the answer, measured as the total time it took to solve the problem.
You're doing work, but what is the goodput in delivering your answer? It's basically the total amount of time it took you to solve the problem.
Something Fails Many Times a Day
At frontier scale a job is synchronous, so a single failed chip among a hundred thousand can stall the whole computation, and at that scale something fails many times a day.
One of them fails it might actually bring the whole thing to a stop right because everyone is counting on everyone else to do their part of the job in order to come up with the answer to a really tough question.
Doubling Every Six Months
Google aims to double the capability to generate tokens every six months, and as much of each doubling tends to come from software and model work as from new silicon.
You have to double the capability of that hardware to generate tokens every six months.
Betting Against the Bitter Lesson
Building a custom accelerator for one workload was contrarian in 2013 when the wisdom was that specialization never wins, but a durable workload made it a massively successful bet that kept generalizing.
So it was a bet. It turned out to be a massively successful bet.
How Far to Specialize
The more a chip specializes, the faster and more power-efficient it gets but the less flexible it is, so the real question is how durable the target workload will be.
this opportunity where the more you specialize to a particular workload the less flexible it is the faster the more power efficient the hardware is going to be. So it is this art and it's this projection of what are you designing to and how persistent is that workload.
Changing the Chip in Flight
Because Google designs both the model and the silicon in the same rooms, DeepMind can intercept a chip weeks before tape-out and change its architecture, something that would be far harder across company lines.
we can intercept and we can make changes to the chip literally the chip architecture in flight which would be somewhere between hard and impossible to do if we were working across company boundaries.
Agents Reshape the Data Center
A long-horizon agent has no human in the loop to rate-limit it, so requests fire in milliseconds instead of seconds and demand for CPUs, networking and storage climbs alongside the accelerators.
there's no human in the loop that is going to naturally rate limits how quickly requests are going to go to the model
Power Is the Binding Constraint
Of all the hard constraints, Vahdat names power as the single most fundamental one, planned years in advance with utilities and sometimes met by generating power locally and feeding it back to the grid at peak.
I would say that power is the single most fundamental constraint that we face