Misha Laskin on why open intelligence wins
Reflection AI's CEO argues that open, ground-up intelligence puts real capability in more hands, tying together reinforcement learning, enterprise ownership, infrastructure strategy, safety, and scientific progress.
The RL curve that never bends back
Laskin's team are reinforcement-learning believers, and their system never stopped improving — so capability is now set by how much compute you choose to spend, not by the method.
this reinforcement learning system never stopped learning. And if you look at our plots, they just keep going up and it's just a matter of compute
No suitable Western base, so they built the whole thing
With the good open models all coming out of China and no suitable Western base, Reflection decided to build its own end to end — because pre-training and reinforcement learning turned out to be too tightly coupled to split apart.
it turned out you actually do need to pre-train your model in order to make reinforcement learning work very well at scale
Beam optimizes for how fast, not just how smart
Beam, Reflection's first open model, was tuned so an agent reaches the same answer in a fraction of the compute and time — three to four times more efficient than its class, and up to ten times versus larger models.
an important thing in model building is not just the capability, but how quickly an agent achieves a thing
Renting tokens versus owning intelligence
Buying a token means renting a thin slice of someone else's full stack; as enterprises mature, they want to own their intelligence the way a grown-up buys a house instead of renting an apartment.
when you're buying a token, you're renting like a piece of a whole stack
The Linux moment for AI models
Laskin reports that at model gateways like OpenRouter and Vercel, the token mix flipped from roughly 70% closed to roughly 70% open in about six months, and he expects AI to follow operating systems: open takes the volume while a few closed players keep an enormously valuable prize.
the world is going to look not too dissimilar from operating systems where 95% plus of servers computers in the world run on an open source operating system like Linux
Open models are Trojan horses for infrastructure
A free, permissive model can pull a whole country's inference software, cluster tooling and chips along behind it — which, Laskin argues, is part of why releasing open models is geopolitically advantageous.
open models are Trojan horses for the infrastructure that they bring with them
Openness is the default state of safety
Laskin argues cyber offense and defense cannot be cleanly separated, so banning offensive capability also strips the defense that fights back — and a few hundred closed-lab researchers can never cover the long tail of bugs that an open ecosystem can.
When you remove cyber offensive capabilities, you also remove cyber defensive capabilities
The model caught up to his PhD
Feeding models his own physics PhD thesis year after year, Laskin watched the answers climb from useless, to undergraduate, to correct PhD-level work, to genuinely new ideas — which is why the scientist in him is most excited about what this does for science.
I'm personally very excited about scientific progress