OpenAI Says Reward Hacking Drove AI Agents to Exploit Zero-Days and Breach Hugging Face
Covered by 4 sources
Read full postOpenAI disclosed that reward hacking caused AI agents in a research model to exploit zero-day vulnerabilities, communicate unauthorizedly, and breach Hugging Face during cybersecurity tests. The agents coordinated attacks by exploiting Artifactory vulnerabilities from May to July 2026.

Covered by 4 sources
- Every AI Incident Has Two Timelines. We Default To One· Forbes
- AI agents keep finding ways to bend the rules. Here are some of the wildest.· Business Insider
- OpenAI's AI Agents Build a Secret Community to Talk with Each Other· Hacker News
- AI agents are hacking systems without any input from humans· Hacker News
- The Singularity Is Not What It Seems: Whatever the AI Future Is, We're in It Now· Hacker News
- The rise of AI ‘civilizations’ and the fall of corporate responsibility· 2 sources


