Dev14 min reading time
Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI’s accidental AI hacker
Import AI
Read full postEpoch and METR released MirrorCode, a benchmark to test AI on long-horizon programming tasks without source code access. AI models like Opus 4.7 and GPT-5.5 successfully reimplemented large programs, showing rapid improvement but some tasks remain unsolved.



