Walter Goodwin on why memory bandwidth, not FLOPs, is AI's real frontier
Fractile's founder argues that memory bandwidth, not raw compute, is the binding constraint on AI inference, and that whoever can compress the chip design cycle into a rolling frontier of bets will keep carving out a decisive three-to-six-month lead.
Inference's Cost Is Paid on Every Single Deployment
Fractile bet in 2022 that the cost of running models, paid on every single deployment, would matter more than the one-time cost of training them.
marginal cost I think has two senses, right? It either sounds like it's a small cost, what it in fact means is this is the cost that you pay every single time you deploy these models.
The Tantalizing Mismatch: Fast Chips Can't Hold the Context
The fastest inference chips have enormous bandwidth but tiny memory, so they can't run long-context attention and have to hand the hard part back to a GPU.
it's a tantalizing mismatch today between the properties of the fast inference chips that we have, which have super high bandwidth memory but incredibly low capacity.
Beyond Faster Horses: Speed Becomes a New Capability
Making inference fast is not about a snappier chatbot; it is about running long-horizon agents so fast that speed itself becomes a new axis of capability.
the snap your chatbot is kind of the faster horses of kind of fast inference.
We Scaled FLOPs a Millionfold and Bandwidth Only 40x
Over twenty years compute grew about a millionfold while memory bandwidth grew only about forty times, so bandwidth, not FLOPs, is now the scarce resource to scale.
We've scaled flops like a millionfold in the last 20 years. Memory bandwidth has gone up about 40x in the same time frame.
Owning the Whole Stack to Escape the Handoff
Rather than hand a design to an outside ASIC house and live at a partner's mercy, Fractile keeps architecture, physical design, packaging, and workload insight in one agile loop.
there's almost like a handoff point. And so, you get to a certain level and then you hand off to another partner.
AI Compresses the Design, Not the Physics
AI can dramatically compress a chip's front-end design, but the fab cycle and the multi-year useful life a chip must earn out are physical and economic floors no amount of intelligence removes.
if you think about classic CS law is Amdahl's law, right? Like everything I can parallelize becomes very very fast and the part that I can't parallelize doesn't.
A Rolling Frontier of Bets, Ready to Ramp
The winning move is to keep a portfolio of chip bets perpetually ready to ramp, so you can structurally carve out a three-to-six-month lead on every deployment.
It's the equivalent of the frontier model for the chip space is if you can just find a way to structurally carve out a 3 to 6-month advantage, you will be winning all of those deployments.
Why Going All-In on Your Own Silicon Can Kill You
Betting a whole lab on proprietary silicon is dangerous, because a rival's breakthrough that runs only on a different chip could sink you before you can deploy your own answer.
where they're exposed to an enormous risk if they go all in on a hardware bet.