Waddle Labs & RoboCurve on Why One LLM Can Run Any Robot
Two frontier startups argue that the fastest path to capable robots is not a robotics-specific model but one strong general-purpose LLM that writes code and calls tools to control any robot, fed with data-rich modalities like coding and computer use.
Skip the Robot Model, Use a General LLM
<strong>Rather than train a robotics-specific model on scarce robot data, point one strong general-purpose LLM at the robot</strong> and let it write the control code.
to me it's like, okay, why not just build a really good foundational LLM and I use that to control robots, rather than training like a model that's specifically for robotics and that's more dependent on just robot data.
RT-2: A Language Model That Outputs Robot Poses
<strong>RT-2 fine-tuned a web-pretrained language model to emit end-effector poses instead of words</strong>, proving that language pretraining transfers to physical control.
One of the earliest successful approaches of using AI on robots is the RT2 paper, where they use a pre-trained language model on web uh text and images and use that to control uh robots.
It's the Data Modality, Not the Architecture
<strong>Since VLAs are already built on language models, the real lever is not the architecture but which data modality you pour in</strong> — and code is a modality models understand well.
So perhaps the bit of lesson here isn't necessarily what architecture you build necessarily, but it's what data is most useful, right?
Code as Policies: One-Shot, No Robot Data
<strong>Given a few Python functions, a coding agent can write control code one-shot — with no extra robot data</strong> — because it is already trained on so much code.
And I think it's this one-shot ability, this in-context exploration ability, that really motivated a lot of later work to continue exploring, including us, to continue exploring how we can apply LLMs to robotics.
The Harness: Compress Experience Into Skills
<strong>In-context learning is cheap but saturates fast, so the harness distills what a robot learns into reusable programs</strong> that future agents can call instead of re-deriving.
writing out these skills and writing out these memories is a form of consolidating. It's it's like a form of distillation um from past experience for your future agents to use.
Why Computer-Use Data Teaches the Physical World
<strong>Computer-use data, including CAD interactions, quietly teaches a model the spatial reasoning robots need</strong> — the same dragging and orbiting skills used on screen can transfer to robot control.
if you feed a computer used data into a big model like Astra, um and by computer used data I mean like you drag a cursor around on a screen to orbit some CAD object in order to design in Blender, this tells you um like how to reason about spaces.
General-Purpose Robots Are About Two Years Out
<strong>Model latency is falling roughly 2x per month, and the founders expect general-purpose robots within two years</strong> — a ChatGPT-style moment they say society is unprepared for.
There's some consensus within the frontier labs and also in the robotics foundation uh robotics companies that we will have general-purpose robots within the next 2 years or even earlier. And this is something that society is probably unaware of or even unprepared for.