Latent Space

Alex Zhang on why the system around the model is the new frontier

Alex Zhang· PhD researcher, MIT at MIT
·~103 min·English·Latent Space
LLMAgentsReasoningTrainingGPU
TL;DR

MIT PhD researcher Alex Zhang, creator of recursive language models (RLMs), argues today's frontier models are already strong — so much of the untapped leverage now lives in the shape of the system around a model: the harness, recursive self-calls, how a model is trained for that harness, and even its output space.

01Core Mental Model

A Language Model That Calls Itself

<strong>An RLM is a harness whose only tool is code</strong> — the model writes programs that keep context in a code environment and can call itself as a sub-agent.

An RLM is basically just a harness design where the only tool in the harness is code.

— Alex Zhang, Latent Space
Key Insight
The radical move here isn't a new neural architecture — it's promoting the Python interpreter to the model's reasoning substrate. Because context sits in the code environment, the model stops being imprisoned by its context window; because the harness can call itself, the same abstraction can stretch from a one-line task toward far longer work — even if reliably sustaining it is still unproven.

02Why Harnesses Exist

Next-Token Prediction Is an Awkward Shape

<strong>Next-token prediction is an awkward form for hard tasks</strong>, so a harness is really an opinionated program that form-fits the model to the problem.

a harness is a very very opinionated program over how you want a language model to be form fit over a problem

— Alex Zhang, Latent Space
Key Insight
If the harness is just encoded opinion, then much of today's agent quality is engineering taste expressed as control flow — portable, forkable, and nowhere near optimized. The awkwardness of next-token prediction is the thing every harness is quietly papering over.

03The Harness Tax

Almost Every Harness Is the Same Loop

<strong>Almost every harness is the same two-calls-in-a-loop</strong> — which is exactly why there is still room to design genuinely different ones.

most harness choices don't matter because all of these harnesses are the same

— Alex Zhang, Latent Space
Key Insight
The 'harness tax' finding is uncomfortable for a whole category of startups selling a thinner wrapper: if the wrappers are interchangeable, the moat has to be a model trained for the harness, or a harness shape nobody else has. Sameness at the bottom is what makes the top worth attacking.

04Compositional Generalization

Keep Every Call In-Distribution

Train an RLM on short tasks and it generalizes to ones 8&ndash;30&times; longer, because <strong>every individual call stays in-distribution even when the whole task is not</strong>.

it turns out that when you take this strategy that they learned, it is directly transferable to the longer length. Like they're effectively the same program.

— Alex Zhang, Latent Space
Key Insight
'Locally in-distribution' reframes generalization as something you engineer, not something you wait for scale to grant. If the harness decides how far a model's training reaches, then data efficiency becomes a design problem — train on fewer environments, and let composition stretch them across a far wider class of tasks.

05Beyond the Decoder

The Output Space Is a Free Variable

<strong>The output space is a tunable variable, not a given</strong> — a model just models language, so the auto-regressive text decoder is one design point among many.

a language model is just modeling language it doesn't have to be this transformer decoder

— Alex Zhang, Latent Space
Key Insight
Fixing the output space is a bet against the labs' monopoly on frontier decoders. You cannot out-scale OpenAI on an auto-regressive decoder — but you can win on a trade-off they have no incentive to explore — using a specialized, fixed-output model when a task only needs a binary answer, instead of paying for unnecessary general-purpose generation.

06Swarms at Scale

Composition Turns a Swarm Into an Answer

<strong>What makes a $40M agent swarm actually solve a problem is composition</strong> — feeding the right information to a model that is already smart enough, not the harness details.

it's very exciting that we even have the option to point $40 million at a problem and solve it

— Alex Zhang, Latent Space
Key Insight
The reported swarm result expands what seems solvable, though its ~$40M cost is only an estimate at public pricing. And Zhang's hunch that most of the swarm is slop reframes the open research question: not whether swarms work, but how to stop paying for the agent work that never touches the answer. Scale made the impossible possible; efficiency is the part still unsolved.

07Research Taste

The Trivial-Looking Idea Is the Bet

<strong>A PhD's real edge is taking big bets on ideas the field is quick to call trivial</strong> — SWEBench, QuietStar and RLMs all drew that reaction before they mattered.

the research is just never going to be that interesting because you kind of need to take big bets if you're going to be in academia

— Alex Zhang, Latent Space
Key Insight
Zhang inverts the usual signal: in a field optimized for legible progress, the dismissive 'why is this useful?' reaction marks the ideas with the least competition. For a resource-poor academic who cannot win on compute, that neglect is the only place left to beat a frontier lab.

08Capability Overhang

It's a Skill Issue, Not a Smarts Issue

<strong>Frontier models still can't do a simple month-long job reliably</strong> — and Zhang argues that gap is a harness problem, not a capability one.

I think that it genuinely is a skill issue of you can get a model to be as good as let's say like just some 18-year-old high school kid doing some job

— Alex Zhang, Latent Space
Key Insight
If the bottleneck is the harness and not the weights, a large share of today's models' economic value is already sitting unclaimed — a capability overhang that a better harness, not a bigger training run, unlocks. Zhang's wager is that 'try harder' is a real engineering program, not a taunt.