Sequoia Capital

Amin Vahdat on Goodput, Not FLOPS

Amin Vahdat· Chief Technologist for AI Infrastructure at Google
·~64 min·English·Sequoia Capital
GPUInferenceTrainingAgentsAI InfrastructureBusiness Strategy
TL;DR

Google aims to double its token-serving capacity every six months, with as much of the gain coming from software as from hardware, and measures itself on goodput under real-world failures rather than peak FLOPS.

01The Metric

Goodput, Not FLOPS

The number that matters is not a chip's theoretical FLOPS but goodput: the useful work that actually reached the answer, measured as the total time it took to solve the problem.

You're doing work, but what is the goodput in delivering your answer? It's basically the total amount of time it took you to solve the problem.

— Amin Vahdat, Sequoia Capital
Key Insight
Choosing goodput as the scorecard quietly moves the industry's goalposts. It rewards the unglamorous work of failure recovery and scheduling rather than the peak-FLOPS figure that looks best on a spec sheet.

02The Failure Reality

Something Fails Many Times a Day

At frontier scale a job is synchronous, so a single failed chip among a hundred thousand can stall the whole computation, and at that scale something fails many times a day.

One of them fails it might actually bring the whole thing to a stop right because everyone is counting on everyone else to do their part of the job in order to come up with the answer to a really tough question.

— Amin Vahdat, Sequoia Capital
Key Insight
If one failure in a hundred thousand parts can stall a synchronous job, then at this scale reliability engineering, not raw chip speed, is what separates two data centers running identical hardware.

03The Pace

Doubling Every Six Months

Google aims to double the capability to generate tokens every six months, and as much of each doubling tends to come from software and model work as from new silicon.

You have to double the capability of that hardware to generate tokens every six months.

— Amin Vahdat, Sequoia Capital
Key Insight
If software contributes as much as silicon to each doubling, a lab's compiler, runtime and model teams are effectively part of its hardware roadmap, and a buyer of chips alone would struggle to match the same tokens per watt.

04The Bet

Betting Against the Bitter Lesson

Building a custom accelerator for one workload was contrarian in 2013 when the wisdom was that specialization never wins, but a durable workload made it a massively successful bet that kept generalizing.

So it was a bet. It turned out to be a massively successful bet.

— Amin Vahdat, Sequoia Capital
Key Insight
The bet paid off partly because Google already owned a workload large and durable enough to amortize a custom chip. That precondition, more than the chip design itself, is what most would-be TPU competitors lack.

05The Dial

How Far to Specialize

The more a chip specializes, the faster and more power-efficient it gets but the less flexible it is, so the real question is how durable the target workload will be.

this opportunity where the more you specialize to a particular workload the less flexible it is the faster the more power efficient the hardware is going to be. So it is this art and it's this projection of what are you designing to and how persistent is that workload.

— Amin Vahdat, Sequoia Capital
Key Insight
Keeping each specialized chip able to do the other's job is a hedge against a six-year forecast being wrong. It gives up some peak efficiency in exchange for the option to rebalance inference and training after the hardware is already deployed.

06Co-Design

Changing the Chip in Flight

Because Google designs both the model and the silicon in the same rooms, DeepMind can intercept a chip weeks before tape-out and change its architecture, something that would be far harder across company lines.

we can intercept and we can make changes to the chip literally the chip architecture in flight which would be somewhere between hard and impossible to do if we were working across company boundaries.

— Amin Vahdat, Sequoia Capital
Key Insight
Owning both the model and the chip turns a normally fixed constraint into something negotiable just before tape-out. Doing the same across company lines is possible but much harder, which is the real edge an integrated model-and-silicon stack buys.

07Shifting Workload

Agents Reshape the Data Center

A long-horizon agent has no human in the loop to rate-limit it, so requests fire in milliseconds instead of seconds and demand for CPUs, networking and storage climbs alongside the accelerators.

there's no human in the loop that is going to naturally rate limits how quickly requests are going to go to the model

— Amin Vahdat, Sequoia Capital
Key Insight
Taking the human out of the loop quietly rebuilds the data center around the CPU and the network again. The accelerator is the headline, but agent workloads make the surrounding orchestration the thing that keeps the expensive chips busy.

08The Ceiling

Power Is the Binding Constraint

Of all the hard constraints, Vahdat names power as the single most fundamental one, planned years in advance with utilities and sometimes met by generating power locally and feeding it back to the grid at peak.

I would say that power is the single most fundamental constraint that we face

— Amin Vahdat, Sequoia Capital
Key Insight
That Google is seriously pursuing data centers in orbit shows how far the energy constraint can push infrastructure choices, even when cooling and in-space repair get harder. It points to energy, not chips, as the scarce resource, without proving orbital is the better answer.