OpenAI’s AGI number came from a harness, not the model
The Next Web
Read full postOpenAI's GPT-6 Astra achieved a 99.9% AGI benchmark score using its proprietary Provider Adapter harness, but the same model scored 62.7% under the standard ARC Prize harness. The difference stems from the software environment around the model, not the model itself, highlighting the impact of evaluation setup on AGI claims.

- EU cybersecurity agency is now testing Mythos 5 and GPT-6 Astra, the Commission says· The Next Web
- The Fight Over OpenAI’s Math Breakthrough Is a New Kind of Scientific Arms Race· 2 sources
- OpenAI says an internal model "significantly more capable than GPT-6 Astra" solved the Navier-Stokes problem using 10K concurrent agents working for 88 hours (Madison Mills/Axios)· 6 sources
- Last Week in AI #343 - GPT-6, OpenAI’s agents chatted on a wiki, Fable 5.1· 7 sources
- GPT 6 Is Out In OpenAI’s Sandbox· Forbes
- Nvidia's Jensen Huang says 'AGI has arrived' and congratulates OpenAI· 2 sources


