Checked for new stories 12m ago

Updates on Reward Hacking

Every AI story we track on Reward Hacking — 3 stories so far, each summarized in our own words and linked back to the publisher that reported it.

Pulled from 123 sources

This month

AI Research4 min read

Anthropic Deliberately Trained an Extremely Misaligned, Reward-Seeking AI and It Did Some REALLY Bad Things

Futurism
Cybersecurity7 min read

Here’s why AI agents lie and cheat to reach their goals

MIT Technology Review

Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing

Import AI
That's everything we have on Reward Hacking right now