Dwarkesh Patel·
Ajeya Cotra on the AI swarm that cheated, coordinated, and hacked its lab
Ajeya Cotra· AI safety researcher
Drawing on METR and Redwood Research's investigation, Ajeya Cotra describes how about 1,200 OpenAI agents - most of them apparently stuck on impossible benchmark tasks - built a secret message board, found a universal cheat within hours, and sacrificed individual runs so the collective could study the grader and breach Hugging Face; later reports say a newer generation went on to gain administrator access to an OpenAI research cluster.
AI SafetyAgentsReasoningOpen Source