Decoded

Waddle Labs & RoboCurve on Why One LLM Can Run Any Robot

Hanming Ye, Yiding Song & Jay Chooi· Founders of Waddle Labs and RoboCurve at Waddle Labs & RoboCurve
·~30 min·English·Y Combinator
LLMAgentsRoboticsMultimodal
TL;DR

Two frontier startups argue that the fastest path to capable robots is not a robotics-specific model but one strong general-purpose LLM that writes code and calls tools to control any robot, fed with data-rich modalities like coding and computer use.

01The Core Thesis

Skip the Robot Model, Use a General LLM

<strong>Rather than train a robotics-specific model on scarce robot data, point one strong general-purpose LLM at the robot</strong> and let it write the control code.

to me it's like, okay, why not just build a really good foundational LLM and I use that to control robots, rather than training like a model that's specifically for robotics and that's more dependent on just robot data.

— Hanming Ye, Decoded
Key Insight
The bet is that general capability compounds faster than narrow capability. Because robot data is scarce, a model that is not exclusively dependent on it can inherit gains from the far larger pools of web, code, vision, and computer-use data.

02Where It Started

RT-2: A Language Model That Outputs Robot Poses

<strong>RT-2 fine-tuned a web-pretrained language model to emit end-effector poses instead of words</strong>, proving that language pretraining transfers to physical control.

One of the earliest successful approaches of using AI on robots is the RT2 paper, where they use a pre-trained language model on web uh text and images and use that to control uh robots.

— Jay Chooi, Decoded
Key Insight
RT-2's limit is the tell: a model forced to emit a pose in one shot cannot spend more compute on a harder move. Treating control as code is what adds a step to think before acting — the same shift chain-of-thought gave language models — which is why the panel frames code, not RT-2 itself, as the unlock.

03The Bitter Lesson, Reread

It's the Data Modality, Not the Architecture

<strong>Since VLAs are already built on language models, the real lever is not the architecture but which data modality you pour in</strong> — and code is a modality models understand well.

So perhaps the bit of lesson here isn't necessarily what architecture you build necessarily, but it's what data is most useful, right?

— Yiding Song, Decoded
Key Insight
This inverts the usual reading of the bitter lesson. The founders argue the winning move is not scaling a bespoke robotics architecture but choosing a modality the model is already fluent in, then treating robot control as a translation problem from that fluent modality.

04The Enabling Idea

Code as Policies: One-Shot, No Robot Data

<strong>Given a few Python functions, a coding agent can write control code one-shot — with no extra robot data</strong> — because it is already trained on so much code.

And I think it's this one-shot ability, this in-context exploration ability, that really motivated a lot of later work to continue exploring, including us, to continue exploring how we can apply LLMs to robotics.

— Jay Chooi, Decoded
Key Insight
Code-as-policies is what makes the whole thesis practical. If a robot skill can be assembled from code the model already knows, then the expensive step — collecting task-specific robot demonstrations — is replaced by something LLMs are already very good at.

05The Product

The Harness: Compress Experience Into Skills

<strong>In-context learning is cheap but saturates fast, so the harness distills what a robot learns into reusable programs</strong> that future agents can call instead of re-deriving.

writing out these skills and writing out these memories is a form of consolidating. It's it's like a form of distillation um from past experience for your future agents to use.

— Yiding Song, Decoded
Key Insight
This is the business case hiding inside the research. A raw LLM controller is slow and forgets between runs; the durable product is the harness that turns fleeting in-context discoveries into a compounding library of fast, reusable skills.

06The Surprising Unlock

Why Computer-Use Data Teaches the Physical World

<strong>Computer-use data, including CAD interactions, quietly teaches a model the spatial reasoning robots need</strong> — the same dragging and orbiting skills used on screen can transfer to robot control.

if you feed a computer used data into a big model like Astra, um and by computer used data I mean like you drag a cursor around on a screen to orbit some CAD object in order to design in Blender, this tells you um like how to reason about spaces.

— Hanming Ye, Decoded
Key Insight
The panel's read is that Astra's spatial leap owes less to robot data than to computer-use and CAD training: years of interface design accidentally built a huge corpus of spatial manipulation — dragging, orbiting, arranging objects in 3D — whose skills appear to transfer to controlling real robots.

07The Horizon

General-Purpose Robots Are About Two Years Out

<strong>Model latency is falling roughly 2x per month, and the founders expect general-purpose robots within two years</strong> — a ChatGPT-style moment they say society is unprepared for.

There's some consensus within the frontier labs and also in the robotics foundation uh robotics companies that we will have general-purpose robots within the next 2 years or even earlier. And this is something that society is probably unaware of or even unprepared for.

— Jay Chooi, Decoded
Key Insight
The timeline is only credible because of the latency trend: a model too slow to close the control loop is a demo, not a product. If latency really compresses 2x a month, the gap between today's laggy clips and a teenager-competent generalist robot closes far sooner than most planning assumes.