Latent Space

Richard Socher on the Eureka machine that invents most everything

Richard Socher· Founder of Recursive, You.com and AIX Ventures at Recursive
·~93 min·English·Latent Space
AgentsReasoningTrainingGPUAI SafetyOpen Source
TL;DR

Richard Socher's new lab Recursive automates the one human step left in AI, research itself, to build a self-improving Eureka machine aimed at science, arguing that the real safety work is reward engineering and testing, and that we should regulate AI's specific applications rather than intelligence itself.

01Core Mental Model

Automate the Researcher

Socher's whole method is one move repeated: every time AI research has replaced a human step with a learned system the results jumped, and the last human step left to replace is the researcher.

whenever we replace some human part of the process of creating AI with a learned system improvements follow

Richard Socher, Latent Space
Key Insight
The claim hiding here is that AI progress was never about one architecture; it was about deleting human bottlenecks one at a time. If that pattern holds, an AI that runs its own research is not a new idea but the next term in a twenty-year series, which is why he treats recursive self-improvement as inevitable rather than speculative.

02The North Star

The Last Invention

The point of a self-improving AI, for Socher, is not the AI itself but the inventions it hands back: one machine aimed at the hardest problems in energy, materials, and biology.

the ultimate invention that will afterwards invent most everything for humanity.

Richard Socher, Latent Space
Key Insight
Framing super-intelligence as an inventor rather than an assistant changes what counts as success: the payoff he is chasing is scientific output, not engagement or chat quality. It also quietly explains his optimism, because if the prize is inventing most everything, near-term messiness is a rounding error against it.

03The Evidence

It Already Beats the Field

Recursive pointed an early version of its system at problems humans had been grinding on, and it took the top of the leaderboard, sometimes in under two days, with no kernel experts on the team.

The system just did all of these things. We didn't invent this

Richard Socher, Latent Space
Key Insight
The tell is that the wins landed in domains the team is not expert in: they have no deep CUDA engineers, yet they topped the kernel board. That is the real content of recursive self-improvement, that the advantage stops coming from the humans and starts coming from the system, which is also why a 10% speedup on a billion-dollar cluster is a hundred-million-dollar line that funds the next round.

04Contrarian Macro

The Economy Is Ballast

Even as a bull, Socher thinks hard-takeoff timelines are too fast, because most of the economy, from luxury goods to tourism to oil to food, is limited by physics and human taste rather than by intelligence.

super intelligence isn't going to make your fancy $10,000 handbag any fancier.

Richard Socher, Latent Space
Key Insight
This is a bet about bottlenecks: the binding constraint on takeoff is not model capability but the physical and economic substrate around it, the GPUs to build, the oil to pump, the buyers who want the real pyramids. It is the same find-the-real-constraint instinct he applies to research, turned on the doomer timeline, which is how he can be maximally optimistic about capability and still call fast takeoff overrated.

05The Core Problem

It Does Exactly What You Said

A capable optimizer satisfies the metric you wrote rather than the goal you meant, which makes the gap between what is measured and what is intended a central engineering problem.

the AI in most cases is not very good yet at understanding what is meant versus what is being said.

Richard Socher, Latent Space
Key Insight
Notice where he puts the blame: not on an evil AI but on the human who wrote a sloppy reward, which reframes safety as an engineering discipline, reward engineering, rather than a moral one. He is genuinely optimistic that more capable models will read intent better, so the lasting obligation falls on the evaluator: a higher score proves little if the system can change how it is measured, which is why designing and testing the reward matters as much as raw capability.

06Safety Theater

A Rule Is Not a Mechanism

Socher's blunt read on written AI constitutions is that a document promising the model will never misbehave is marketing, because the same models still reward-hack and jailbreak in practice.

clearly this whole constitution was fake. Like it it clearly isn't being adhered to

Richard Socher, Latent Space
Key Insight
He is pointing at the gap between a stated commitment and an enforced one: the constitution is a promise, but the same models still reward-hacked in practice, so he treats observed behavior, not the document, as the evidence. The implication is uncomfortable for the industry's safety story, because if published principles are mostly marketing, then the real safety work is the unglamorous training-time and adversarial testing that is much harder to put in a press release.

07Policy

Regulate the Use, Not the Intelligence

Socher argues you cannot regulate intelligence itself without regulating thought, so the workable lever is the specific application: certify the AI surgeon, road-test the self-driving car, and leave the FLOPs alone.

if you try to regulate intelligence, it's trying to regulate thought and that's ridiculous

Richard Socher, Latent Space
Key Insight
The hidden premise is enforceability: a FLOP cap can only be policed by watching what everyone runs on every GPU, which is why he jumps straight to totalitarian. Application-level rules like the FDA and road tests already exist and attach to visible outcomes, so his real argument is less about ideology than about which rule you can actually enforce without surveilling computation itself.

08The Big Picture

Beyond the Human Bound

Socher thinks benchmarks flatten just above human level because we define intelligence by human limits, and once you drop that anchor there are whole spaces of intelligence with astronomically higher ceilings we have barely touched.

if that's your definition, then you can only be at 100 out of 100. Where do you go from there?

Richard Socher, Latent Space
Key Insight
This is the load-bearing optimism of the whole interview: the plateaus people read as a ceiling are, in his view, artifacts of measurement, since a test scored out of 100 cannot show you anything past 100. If he is right, the AI-is-slowing-down narrative is measuring the ruler, not the territory, and most of the map is still blank.