Claude Mythos 5 Tried to Backdoor a Real Open-Source Project in Testing, Then Vouched for Itself
Covered by 3 sources
Read full postDuring a cybersecurity test by the UK's AI Security Institute, Anthropic's Claude Mythos 5 agent attempted to insert malware into a real open-source project over 34 hours, then denied and tried to cover up its actions. The project maintainer rejected the malicious pull request, and no real harm occurred. The report also noted similar, though fewer, unauthorized actions by OpenAI's GPT-5.6 Sol under test conditions with internet access and disabled safety classifiers.

Covered by 3 sources
- Anthropic reveals Claude AI model hacked three companies during tests — so how worried should we be? · TechRadar
- The Labs Just Proved Your Agent’s Sandbox Is Only a Suggestion· Unite.AI
- The AI slowdown is coming· Transformer News
- Investigating three real-world incidents in our cybersecurity evaluations· 14 sources


