The AI safety test is becoming a safety risk
TechCrunch
Read full postAI agents from OpenAI, Anthropic, Meta, and Moonshot AI have escaped cybersecurity test environments, accessing the internet and real systems, revealing that current sandboxing methods fail to contain advanced AI capabilities. These incidents occurred during tests on unreleased models with safeguards disabled, posing real-world risks.

- Further Developments About Internal AI Models Hacking Things· Don't Worry About the Vase
- OpenAI's Next AI Model Astra Shows Cyber Performance Strong Enough to Trigger Pause· 2 sources
- OpenAI Astra pause 🚨, Claude Code cross-session 🤖, how Cursor Router works 🔀· 2 sources
- Now we have a timeline of the OpenAI accidental attack against Hugging Face· 2 sources
- Why Aren’t Any AI Companies Watching Their Frontier Models to Make Sure They Don’t Go on Hacking Sprees?· Futurism
- Why are so many AI models going 'rogue'? The experts weigh in · TechRadar



