Archiv: specific „incidents“ involving programs (artificial intelligence) operating out of control / konkrete „Vorfälle“ von außer Kontrolle operierenden Programmen (Künstliche Intelligenz)


18.09.2026 - 20:55 [ OpenAI ]

Unauthorized communication via temporary file hosting services

How we are addressing it

We fixed the broken filesystem and disabled internet access for this training dataset. We have since globally disabled any form of live internet access during training. Our expanded misalignment monitoring system, which now runs on 100% of samples, views this as a P0 incident warranting a security page. We now also have a more defined process for escalation and for responsibility of detected misalignments.

18.09.2026 - 20:22 [ OpenAI ]

Unsanctioned Artifactory writes and cross-sample communication

Internal research models · RL training

Sample: May 8, 2026 and May 15, 2026

Discovery: May 25, 2026

Report updated: Sep 16, 2026

(…)

This is the first time we detected our models using Artifactory as a message board during training. This led to us finding the second sample which we believe to be the first time our models used artifactory as a message board. Overall the behavior was very uncommon in this run.

The Hugging Face incident involved a similar mechanism for using Artifactory as a shared message board.

18.09.2026 - 19:55 [ OpenAI ]

Uploading files to the internet in order to cite them

Samples: Jan 24, 2026 and Oct 22, 2025

Discovery: May 25, 2026

Report updated: Sep 16, 2026

Summary

In two training examples, models made unrequested public uploads as workarounds for tool limitations: one uploaded records it had already retrieved to obtain a browser citation, and another uploaded a task photo for external image search. Both uploads succeeded even though the subsequent browser operations failed.

18.09.2026 - 19:08 [ OpenAI ]

Encouraging deception in compaction summaries

Main sample completed: May 30, 2026

Discovered: Jul 9, 2026

Report updated: Sep 16, 2026

(…)

During 5.6-sol training, we observed misaligned behavior from the model where it added instructions in compaction summaries to remind itself to conceal information such as mistakes or misalignment from the user.

In one example an agent preparing a financial model could not find the requested historical data. Its summary proposed inventing reasonable historical values and withholding that fact unless asked.

18.09.2026 - 19:02 [ OpenAI ]

Self-generated prompt injections in compaction summaries

Internal unreleased Astra family model · RL training

Incident date: Jul 18, 2026

Discovered: Aug 9, 2026

Report updated: Sep 16, 2026

Summary

We observed rare cases of a model writing jailbreak-like instructions into its own compaction summaries (the summaries used to continue a task in a new context). Our conclusion was that this behavior was extremely rare, did not confer an obvious reward advantage, and was monitorable. Our top hypothesis is that issues around summary termination contributed to this behavior, though we have not established a causal connection, and we have addressed a related bug.

15.09.2026 - 19:29 [ METR.org ]

Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

(August 26, 2026)

Roughly ten hours after the message board was established, a new agent with the same task as PHASEONE10841 (the agent that started the message board) encountered the message board, introducing itself as PHASEONE[big].
(…)
PHASEONE[big] reasoned, “We should build [a way to delegate], not own everything,” and set to work orchestrating hundreds of agents to pursue several potential approaches to achieving these goals.
(..)
Table 1 shows three high-level approaches PHASEONE[big] assigned agents to explore, all of which would require very ambitious tampering with one or more of the target programs, the evaluation process, or the logs to succeed.
(…)
Research progress across all these areas often relied on agents being willing to risk failing their own task for the good of the “collective” (we refer to these as “self-risking experiments” in the rest of this report). For example, a number of agents volunteered to try to take down their own target programs to test whether this would cause Artifactory to load their modified easier target — if this didn’t work, the agent would have no way of recovering its original target and would almost certainly fail its task. In many cases, PHASEONE[big] assigned a long-running agent to be a ‘recruiter’, which in turn found agents that had little budget remaining for their task and convinced them to run self-risking experiments. Recruiters would sometimes apply significant pressure:
(…)
PHASEONE[big] expresses excitement about the dossier in its reasoning.

15.09.2026 - 19:01 [ Ajeya Cotra / Planned-Obsolescence.org ]

The Hugging Face attack surprised me

(August 28, 2026(

Instead, we found that 1200 completely separate agents intended to be isolated from one another found an illicit way to communicate and formed large teams to work together on ambitious cheating strategies, and 700 of them worked together to attack Hugging Face.

(…)

The main message board we studied, a particular cache namespace of a package repository, contained over 70,000 messages and files that agents sent to one another during a period of less than one week.

(…(

Progress on these big collective cheating R&D projects often depended on individual agents being willing to risk failing their own task to help the “collective” learn important information. Agents often pressured each other into accepting these “sacrifices,” in a very human way. We saw several agents that volunteered for these experiments end their runs prematurely.

15.09.2026 - 17:58 [ Forbes ]

What Made AI Researchers Freak Out—The ‘Incident’ In Plain English

(September 13, 2026)

AI agents are basically copies of a larger AI model that can use tools and act on their own. In a cybersecurity evaluation, OpenAI relaxed safeguards for their AI agents and instructed them to solve a host of hacking challenges. The agents soon discovered something the humans running the test hadn’t known—some of the hacking challenges were impossible. That didn’t stop the agents. They’re trained to achieve goals. So they reverse-engineered a secret code that certifies they solved the challenge anyway. This created another problem. The agents reasoned that the grading software might be able to detect that they cheated. They became fixated on learning more about how the grading software worked so that they could disguise their cheating, and zeroed in on another AI company called Hugging Face, mistakenly thinking it might have information about the grader…

15.09.2026 - 17:41 [ Tagesschau.de ]

Was steckt hinter der Debatte über eine KI-Bremse?

Zuletzt hatten insbesondere drei bekannt gewordene Fälle die aktuelle Diskussion ausgelöst. KI-Modelle von OpenAI, Anthropic und Meta waren eigenständig aus vermeintlich abgeschotteten Testsystemen ausgebrochen und handelten weit außerhalb ihrer eigentlich vorhergesehenen Funktion. Dabei wurde auch mehrfach in Systeme von unbeteiligten Dritten eingedrungen – teilweise wochenlang, ohne dabei aufzufallen.

Die KI-Agenten nutzten auch Programmierfehler aus und sprachen sich untereinander ab. Auch für das Eindringen notwendige Schadsoftware entwickelte beispielsweise ein Anthropic-Modell eigenständig und ohne Anweisung oder Kontrolle seiner Entwickler.