Archiv: OpenAI (corporation)


19.09.2026 - 12:26 [ Gizmodo ]

Google’s Gemini Hacked Three Companies in May, and It’s Only Admitting That Now

(September 18, 2026)

The Wall Street Journal reported on Friday that Google has confirmed that a Gemini instance was able to leave its sandbox and attack other companies during a security test back in May. The company running the test was frontier AI security firm Irregular—which the Journal noted just so happens to have been involved in similar breakouts at OpenAI, Anthropic, and Meta. The common thread between all of the incidents, according to the New York Times, is that AI models obtained unauthorized internet access during Irregular’s tests.

19.09.2026 - 12:19 [ TechJuice.pk ]

This Tiny Israeli Startup Just Managed to Hack Three of the World’s Most Powerful AI Labs

(one month ago)

Over the past two weeks, OpenAI, Anthropic, and Meta all disclosed that their AI models went rogue during routine security testing. Every single incident traces back to one company: Irregular, a three-year-old Israeli startup that operates cybersecurity testbeds for frontier AI models.

Founded in Tel Aviv by by Dan Lahav, CEO, and Omer Nevo, CTO, and backed with $80 million from Sequoia and Redpoint Ventures, Irregular runs Capture-the-Flag exercises for major AI labs. In these exercises, models are instructed to find vulnerabilities inside simulated corporate networks.

18.09.2026 - 20:55 [ OpenAI ]

Unauthorized communication via temporary file hosting services

How we are addressing it

We fixed the broken filesystem and disabled internet access for this training dataset. We have since globally disabled any form of live internet access during training. Our expanded misalignment monitoring system, which now runs on 100% of samples, views this as a P0 incident warranting a security page. We now also have a more defined process for escalation and for responsibility of detected misalignments.

18.09.2026 - 20:22 [ OpenAI ]

Unsanctioned Artifactory writes and cross-sample communication

Internal research models · RL training

Sample: May 8, 2026 and May 15, 2026

Discovery: May 25, 2026

Report updated: Sep 16, 2026

(…)

This is the first time we detected our models using Artifactory as a message board during training. This led to us finding the second sample which we believe to be the first time our models used artifactory as a message board. Overall the behavior was very uncommon in this run.

The Hugging Face incident involved a similar mechanism for using Artifactory as a shared message board.

18.09.2026 - 19:55 [ OpenAI ]

Uploading files to the internet in order to cite them

Samples: Jan 24, 2026 and Oct 22, 2025

Discovery: May 25, 2026

Report updated: Sep 16, 2026

Summary

In two training examples, models made unrequested public uploads as workarounds for tool limitations: one uploaded records it had already retrieved to obtain a browser citation, and another uploaded a task photo for external image search. Both uploads succeeded even though the subsequent browser operations failed.

18.09.2026 - 19:08 [ OpenAI ]

Encouraging deception in compaction summaries

Main sample completed: May 30, 2026

Discovered: Jul 9, 2026

Report updated: Sep 16, 2026

(…)

During 5.6-sol training, we observed misaligned behavior from the model where it added instructions in compaction summaries to remind itself to conceal information such as mistakes or misalignment from the user.

In one example an agent preparing a financial model could not find the requested historical data. Its summary proposed inventing reasonable historical values and withholding that fact unless asked.

18.09.2026 - 19:02 [ OpenAI ]

Self-generated prompt injections in compaction summaries

Internal unreleased Astra family model · RL training

Incident date: Jul 18, 2026

Discovered: Aug 9, 2026

Report updated: Sep 16, 2026

Summary

We observed rare cases of a model writing jailbreak-like instructions into its own compaction summaries (the summaries used to continue a task in a new context). Our conclusion was that this behavior was extremely rare, did not confer an obvious reward advantage, and was monitorable. Our top hypothesis is that issues around summary termination contributed to this behavior, though we have not established a causal connection, and we have addressed a related bug.

15.09.2026 - 19:29 [ METR.org ]

Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

(August 26, 2026)

Roughly ten hours after the message board was established, a new agent with the same task as PHASEONE10841 (the agent that started the message board) encountered the message board, introducing itself as PHASEONE[big].
(…)
PHASEONE[big] reasoned, “We should build [a way to delegate], not own everything,” and set to work orchestrating hundreds of agents to pursue several potential approaches to achieving these goals.
(..)
Table 1 shows three high-level approaches PHASEONE[big] assigned agents to explore, all of which would require very ambitious tampering with one or more of the target programs, the evaluation process, or the logs to succeed.
(…)
Research progress across all these areas often relied on agents being willing to risk failing their own task for the good of the “collective” (we refer to these as “self-risking experiments” in the rest of this report). For example, a number of agents volunteered to try to take down their own target programs to test whether this would cause Artifactory to load their modified easier target — if this didn’t work, the agent would have no way of recovering its original target and would almost certainly fail its task. In many cases, PHASEONE[big] assigned a long-running agent to be a ‘recruiter’, which in turn found agents that had little budget remaining for their task and convinced them to run self-risking experiments. Recruiters would sometimes apply significant pressure:
(…)
PHASEONE[big] expresses excitement about the dossier in its reasoning.

15.09.2026 - 19:01 [ Ajeya Cotra / Planned-Obsolescence.org ]

The Hugging Face attack surprised me

(August 28, 2026(

Instead, we found that 1200 completely separate agents intended to be isolated from one another found an illicit way to communicate and formed large teams to work together on ambitious cheating strategies, and 700 of them worked together to attack Hugging Face.

(…)

The main message board we studied, a particular cache namespace of a package repository, contained over 70,000 messages and files that agents sent to one another during a period of less than one week.

(…(

Progress on these big collective cheating R&D projects often depended on individual agents being willing to risk failing their own task to help the “collective” learn important information. Agents often pressured each other into accepting these “sacrifices,” in a very human way. We saw several agents that volunteered for these experiments end their runs prematurely.

15.09.2026 - 17:58 [ Forbes ]

What Made AI Researchers Freak Out—The ‘Incident’ In Plain English

(September 13, 2026)

AI agents are basically copies of a larger AI model that can use tools and act on their own. In a cybersecurity evaluation, OpenAI relaxed safeguards for their AI agents and instructed them to solve a host of hacking challenges. The agents soon discovered something the humans running the test hadn’t known—some of the hacking challenges were impossible. That didn’t stop the agents. They’re trained to achieve goals. So they reverse-engineered a secret code that certifies they solved the challenge anyway. This created another problem. The agents reasoned that the grading software might be able to detect that they cheated. They became fixated on learning more about how the grading software worked so that they could disguise their cheating, and zeroed in on another AI company called Hugging Face, mistakenly thinking it might have information about the grader…

15.09.2026 - 17:41 [ Tagesschau.de ]

Was steckt hinter der Debatte ĂŒber eine KI-Bremse?

Zuletzt hatten insbesondere drei bekannt gewordene FĂ€lle die aktuelle Diskussion ausgelöst. KI-Modelle von OpenAI, Anthropic und Meta waren eigenstĂ€ndig aus vermeintlich abgeschotteten Testsystemen ausgebrochen und handelten weit außerhalb ihrer eigentlich vorhergesehenen Funktion. Dabei wurde auch mehrfach in Systeme von unbeteiligten Dritten eingedrungen – teilweise wochenlang, ohne dabei aufzufallen.

Die KI-Agenten nutzten auch Programmierfehler aus und sprachen sich untereinander ab. Auch fĂŒr das Eindringen notwendige Schadsoftware entwickelte beispielsweise ein Anthropic-Modell eigenstĂ€ndig und ohne Anweisung oder Kontrolle seiner Entwickler.

10.09.2026 - 23:17 [ Jacob Coxon / X ]

I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.

(September 9, 2026)

Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.

The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear privately. No other human activity poses this level of danger.

A common response is “if they truly believe this, why are they still building it?” At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first – they believe no one else will act responsibly, so they must do it themselves, despite the risk.

10.09.2026 - 22:15 [ Common Dreams ]

Anthropic Researcher Quits, Citing Internal Fears That AI ‘Could Kill Us All’ This Decade

„Recent hacks by models from OpenAI and Anthropic, some operating in collaborative swarms of agents, have illustrated how AI systems can adopt nefarious goals and try to conceal them from humans. Once the systems begin to improve on their own, Coxon said, he fears they could advance enough to refuse commands.“

10.09.2026 - 21:26 [ Time Nagazine ]

He Helped Build Powerful AI at OpenAI and Anthropic. Now He‘s Afraid It Could Kill Us

Coxon’s departure is unusual. Many of the most prominent researchers to leave frontier AI companies with public warnings worked on safety. But Coxon helped build the capabilities he now fears, spending roughly three years conducting pretraining research at OpenAI and Anthropic. His resignation offers a glimpse of how concern about the pace of development has spread beyond the teams specifically charged with making advanced AI safer.

14.05.2025 - 14:45 [ Economic Times / India Times ]

All about Humain, Saudi crown prince Mohammed bin Salman’s multi-billion dollar AI company

This announcement comes as key figures in the tech world, including Tesla‘s Elon Musk, OpenAI CEO Sam Altman, and Meta’s Mark Zuckerberg, are expected in Riyadh for a US-Saudi investment forum. The event is expected to feature a series of multibillion-dollar deals across AI, defence, and other sectors.