Rich Sutton & Khurram Javed on why AI must never stop learning
Reinforcement-learning pioneer Rich Sutton and his Oak co-founder Khurram Javed argue that today's LLMs stop learning the moment training ends, and that real intelligence requires continual learning from an infinitely complex world.
Not Him, the Field Is Weird
<strong>All learning is continual</strong> — treating it as a special phase you can switch off is the anomaly, not the other way around.
I'm not weird. The field is weird. The field they need to call it continual learning. It's just learning.
The Bitter Lesson in 26 Words
<strong>Bet on methods that scale with computation</strong> — search and learning — not on hand-coded human knowledge that stops paying off.
don't be distracted by human knowledge as AI traditionally has been many times. Instead focus on learning methods that will scale with computation like search and like learning.
Proof and Casualty at Once
LLMs are <strong>both the Bitter Lesson's best proof and its coming casualty</strong>: they scaled by drinking the internet, then hit the ceiling of a finite internet.
And uh the world is big and the world is massively bigger than everything we stored on the internet.
The World Is Too Big to Simulate
<strong>The world is infinitely more complex than any model of it</strong>, so no amount of synthetic data can replace learning from real experience.
The world is infinitely complex, and any simulation of it is like microscopic.
Frozen the Moment Training Ends
A deployed LLM <strong>never changes its weights</strong>, so it cannot learn anything new while you use it — the opposite of a plastic brain.
Their weights never change.
Catastrophic Forgetting Is Curable
Catastrophic forgetting is <strong>an algorithmic gap, not a law of nature</strong>: per-weight step sizes plus generate-and-test let a network keep learning without erasing what it knew.
Most use cases you have a single stream of data. And then if you apply it to the naive thing, it just completely destroys your prior knowledge in a very um destructive way.
Step 2 Unlocks Everything
<strong>Continual deep learning is step two of the Alberta Plan</strong> because once a system can keep updating its world model, abstraction and planning follow.
And we think that one is like almost the most important because it unlocks everything else. If you could do continual deep learning, you could then continually update your model of the world.
A Quarter of Intelligence
Language is <strong>only about a quarter of intelligence</strong>, so LLMs are a genuine breakthrough that was mistaken for the whole of AI.
All of intelligence is not fluid, capable use of language. There's so much more. It's an important part. You know, it's like 20% or a quarter of intelligence.