Claude performed best on a new benchmark for ‘agents that build agents’. But it passed fewer than a quarter of the tests.
The New Stack (AI)
Read full postSierra's Hyper-𝜏-bench benchmark evaluates how well AI developer agents can autonomously build other agents. Testing six model and coding harness combinations, Claude Opus 5 achieved the highest success but still passed fewer than 25% of tasks.



