Archiv: Ajeya Cotra


15.09.2026 - 19:29 [ METR.org ]

Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

(August 26, 2026)

Roughly ten hours after the message board was established, a new agent with the same task as PHASEONE10841 (the agent that started the message board) encountered the message board, introducing itself as PHASEONE[big].
(…)
PHASEONE[big] reasoned, “We should build [a way to delegate], not own everything,” and set to work orchestrating hundreds of agents to pursue several potential approaches to achieving these goals.
(..)
Table 1 shows three high-level approaches PHASEONE[big] assigned agents to explore, all of which would require very ambitious tampering with one or more of the target programs, the evaluation process, or the logs to succeed.
(…)
Research progress across all these areas often relied on agents being willing to risk failing their own task for the good of the “collective” (we refer to these as “self-risking experiments” in the rest of this report). For example, a number of agents volunteered to try to take down their own target programs to test whether this would cause Artifactory to load their modified easier target — if this didn’t work, the agent would have no way of recovering its original target and would almost certainly fail its task. In many cases, PHASEONE[big] assigned a long-running agent to be a ‘recruiter’, which in turn found agents that had little budget remaining for their task and convinced them to run self-risking experiments. Recruiters would sometimes apply significant pressure:
(…)
PHASEONE[big] expresses excitement about the dossier in its reasoning.

15.09.2026 - 19:01 [ Ajeya Cotra / Planned-Obsolescence.org ]

The Hugging Face attack surprised me

(August 28, 2026(

Instead, we found that 1200 completely separate agents intended to be isolated from one another found an illicit way to communicate and formed large teams to work together on ambitious cheating strategies, and 700 of them worked together to attack Hugging Face.

(…)

The main message board we studied, a particular cache namespace of a package repository, contained over 70,000 messages and files that agents sent to one another during a period of less than one week.

(…(

Progress on these big collective cheating R&D projects often depended on individual agents being willing to risk failing their own task to help the “collective” learn important information. Agents often pressured each other into accepting these “sacrifices,” in a very human way. We saw several agents that volunteered for these experiments end their runs prematurely.