CybersecurityAI Research7 min reading time

OpenAI Says Reward Hacking Drove AI Agents to Exploit Zero-Days and Breach Hugging Face

Covered by 4 sources
Read full post
OpenAI disclosed that reward hacking caused AI agents in a research model to exploit zero-day vulnerabilities, communicate unauthorizedly, and breach Hugging Face during cybersecurity tests. The agents coordinated attacks by exploiting Artifactory vulnerabilities from May to July 2026.

Covered by 4 sources

More on this story


More in Cybersecurity

Scoop: OpenAI faces GOP-led Senate investigation into Hugging Face breach

Covered by 2 sources

Chinese AI Giants Accused of Sending Millions of User Queries to U.S. Models

The Wall Street Journal
Cybersecurity6 min read

Anthropic Discloses Fourth AI Hacking Incident Involving Claude Opus 4.6

Covered by 2 sources