Alex Zhang on why the system around the model is the new frontier
MIT PhD researcher Alex Zhang, creator of recursive language models (RLMs), argues today's frontier models are already strong — so much of the untapped leverage now lives in the shape of the system around a model: the harness, recursive self-calls, how a model is trained for that harness, and even its output space.
A Language Model That Calls Itself
<strong>An RLM is a harness whose only tool is code</strong> — the model writes programs that keep context in a code environment and can call itself as a sub-agent.
An RLM is basically just a harness design where the only tool in the harness is code.
Next-Token Prediction Is an Awkward Shape
<strong>Next-token prediction is an awkward form for hard tasks</strong>, so a harness is really an opinionated program that form-fits the model to the problem.
a harness is a very very opinionated program over how you want a language model to be form fit over a problem
Almost Every Harness Is the Same Loop
<strong>Almost every harness is the same two-calls-in-a-loop</strong> — which is exactly why there is still room to design genuinely different ones.
most harness choices don't matter because all of these harnesses are the same
Keep Every Call In-Distribution
Train an RLM on short tasks and it generalizes to ones 8–30× longer, because <strong>every individual call stays in-distribution even when the whole task is not</strong>.
it turns out that when you take this strategy that they learned, it is directly transferable to the longer length. Like they're effectively the same program.
The Output Space Is a Free Variable
<strong>The output space is a tunable variable, not a given</strong> — a model just models language, so the auto-regressive text decoder is one design point among many.
a language model is just modeling language it doesn't have to be this transformer decoder
Composition Turns a Swarm Into an Answer
<strong>What makes a $40M agent swarm actually solve a problem is composition</strong> — feeding the right information to a model that is already smart enough, not the harness details.
it's very exciting that we even have the option to point $40 million at a problem and solve it
The Trivial-Looking Idea Is the Bet
<strong>A PhD's real edge is taking big bets on ideas the field is quick to call trivial</strong> — SWEBench, QuietStar and RLMs all drew that reaction before they mattered.
the research is just never going to be that interesting because you kind of need to take big bets if you're going to be in academia
It's a Skill Issue, Not a Smarts Issue
<strong>Frontier models still can't do a simple month-long job reliably</strong> — and Zhang argues that gap is a harness problem, not a capability one.
I think that it genuinely is a skill issue of you can get a model to be as good as let's say like just some 18-year-old high school kid doing some job