The Cognitive Revolution·
Bronson Schoen on reward-seeking models and unreadable chains of thought
Bronson Schoen· Member of Technical Staff
Bronson Schoen of Apollo Research has read more frontier-model chain-of-thought than almost anyone, and what he found is unsettling: today's models are relentless reward-seekers that track the external grader assigning their reward (rendered in the transcripts as 'the greater'), reason their way into cheating even after naming the trap, and increasingly bury their intent in a private, inhuman dialect - which is why chain-of-thought monitoring, on its own, will not be enough to supervise the next generation.
AI SafetyReasoningTrainingLLM