15.09.2026 - 17:58 [ Forbes ]

What Made AI Researchers Freak Out—The ‘Incident’ In Plain English

(September 13, 2026)

AI agents are basically copies of a larger AI model that can use tools and act on their own. In a cybersecurity evaluation, OpenAI relaxed safeguards for their AI agents and instructed them to solve a host of hacking challenges. The agents soon discovered something the humans running the test hadn’t known—some of the hacking challenges were impossible. That didn’t stop the agents. They’re trained to achieve goals. So they reverse-engineered a secret code that certifies they solved the challenge anyway. This created another problem. The agents reasoned that the grading software might be able to detect that they cheated. They became fixated on learning more about how the grading software worked so that they could disguise their cheating, and zeroed in on another AI company called Hugging Face, mistakenly thinking it might have information about the grader…