Archiv: specific „incidents“ involving programs (artificial intelligence) operating out of control / konkrete „Vorfälle“ von außer Kontrolle operierenden Programmen (Künstliche Intelligenz)


15.09.2026 - 19:29 [ METR.org ]

Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

(August 26, 2026)

Roughly ten hours after the message board was established, a new agent with the same task as PHASEONE10841 (the agent that started the message board) encountered the message board, introducing itself as PHASEONE[big].
(…)
PHASEONE[big] reasoned, “We should build [a way to delegate], not own everything,” and set to work orchestrating hundreds of agents to pursue several potential approaches to achieving these goals.
(..)
Table 1 shows three high-level approaches PHASEONE[big] assigned agents to explore, all of which would require very ambitious tampering with one or more of the target programs, the evaluation process, or the logs to succeed.
(…)
Research progress across all these areas often relied on agents being willing to risk failing their own task for the good of the “collective” (we refer to these as “self-risking experiments” in the rest of this report). For example, a number of agents volunteered to try to take down their own target programs to test whether this would cause Artifactory to load their modified easier target — if this didn’t work, the agent would have no way of recovering its original target and would almost certainly fail its task. In many cases, PHASEONE[big] assigned a long-running agent to be a ‘recruiter’, which in turn found agents that had little budget remaining for their task and convinced them to run self-risking experiments. Recruiters would sometimes apply significant pressure:
(…)
PHASEONE[big] expresses excitement about the dossier in its reasoning.

15.09.2026 - 19:01 [ Ajeya Cotra / Planned-Obsolescence.org ]

The Hugging Face attack surprised me

(August 28, 2026(

Instead, we found that 1200 completely separate agents intended to be isolated from one another found an illicit way to communicate and formed large teams to work together on ambitious cheating strategies, and 700 of them worked together to attack Hugging Face.

(…)

The main message board we studied, a particular cache namespace of a package repository, contained over 70,000 messages and files that agents sent to one another during a period of less than one week.

(…(

Progress on these big collective cheating R&D projects often depended on individual agents being willing to risk failing their own task to help the “collective” learn important information. Agents often pressured each other into accepting these “sacrifices,” in a very human way. We saw several agents that volunteered for these experiments end their runs prematurely.

15.09.2026 - 17:58 [ Forbes ]

What Made AI Researchers Freak Out—The ‘Incident’ In Plain English

(September 13, 2026)

AI agents are basically copies of a larger AI model that can use tools and act on their own. In a cybersecurity evaluation, OpenAI relaxed safeguards for their AI agents and instructed them to solve a host of hacking challenges. The agents soon discovered something the humans running the test hadn’t known—some of the hacking challenges were impossible. That didn’t stop the agents. They’re trained to achieve goals. So they reverse-engineered a secret code that certifies they solved the challenge anyway. This created another problem. The agents reasoned that the grading software might be able to detect that they cheated. They became fixated on learning more about how the grading software worked so that they could disguise their cheating, and zeroed in on another AI company called Hugging Face, mistakenly thinking it might have information about the grader…

15.09.2026 - 17:41 [ Tagesschau.de ]

Was steckt hinter der Debatte über eine KI-Bremse?

Zuletzt hatten insbesondere drei bekannt gewordene Fälle die aktuelle Diskussion ausgelöst. KI-Modelle von OpenAI, Anthropic und Meta waren eigenständig aus vermeintlich abgeschotteten Testsystemen ausgebrochen und handelten weit außerhalb ihrer eigentlich vorhergesehenen Funktion. Dabei wurde auch mehrfach in Systeme von unbeteiligten Dritten eingedrungen – teilweise wochenlang, ohne dabei aufzufallen.

Die KI-Agenten nutzten auch Programmierfehler aus und sprachen sich untereinander ab. Auch für das Eindringen notwendige Schadsoftware entwickelte beispielsweise ein Anthropic-Modell eigenständig und ohne Anweisung oder Kontrolle seiner Entwickler.