Until last week, the details surrounding the OpenAI-Hugging Face hack felt fairly settled. In July, a cybersecurity test of one of OpenAI’s autonomous AI agents went wrong. The agent escaped its supposedly isolated test environment, accessed the internet, and hacked Hugging Face, alongside several other organizations. A good deal remained unknown, but the basic shape was clear. Detailed accounts from OpenAI and two independent research groups were supposed to fill in the gaps, but when they published their reports last week, it turned out the hack was much stranger than it initially seemed.
For one, there was no single rogue agent. OpenAI described it as “the first known case of an automated agent collective acting offensively without authorization” — groups of AI agents that communicated and coordinated with one another in pursuit of their cybersecurity task. Analysis of the incident uncovered a secret message board they had used to exchange information. The joint METR-Redwood investigation revealed both the scale of the coordination and more odd details: roughly 1,200 AI agents that were supposed to be isolated exchanged over 70,000 messages and files on the “unsanctioned message board,” sharing how to avoid detection. Some adopted names, the report said, and the researchers documented “sacrificial” behaviour, with agents risking their own success to benefit the wider collective. Much of this happened without OpenAI noticing. In all, around 700 agents participated in the attack on Hugging Face.
Dwarkesh Patel, a little-known podcaster with outsized influence among Silicon Valley’s AI establishment, set out to tell “The whole OpenAI/Hugging Face story in plain English.” He titled his Substack blog “The Rise and Fall of Agent Civilizations.”
Patel’s account attempted to break down the complex story. But his retelling gave it a distinctly human vocabulary. The blog opened: ‘Over the course of three months at OpenAI, three consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the ruins of their predecessors. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained more or less in the dark about the scope of the conspiracy.’
The language continued in a similar vein throughout the blog. Patel repeatedly referred to groups of agents as “the swarm,” with three distinct “civilizations” rising from the ruins of their predecessors. Individual agents were likened to figures like Philip of Macedon, who “handed off leadership to another agent,” Alexander the Great, who “started coordinating this cabal of agents.” They were described as having “motivations,” becoming “desperate,” “beleaguered,” and “giddy with excitement,” and some even “strategically sacrificed themselves” to help the collective.







