Dev14 min reading time

Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI’s accidental AI hacker

Import AI
Read full post
Epoch and METR released MirrorCode, a benchmark to test AI on long-horizon programming tasks without source code access. AI models like Opus 4.7 and GPT-5.5 successfully reimplemented large programs, showing rapid improvement but some tasks remain unsolved.

More in Dev

Dev6 min read

How Credit Genie keeps codebase docs fresh with OpenWiki

LangChain
Dev19 min read

Article: When Spec-Driven Development Pays Off

InfoQ (AI, ML & Data)
Dev19 min read

Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

AWS Blog