AgentsMachine Learning9 min reading time

Claude performed best on a new benchmark for ‘agents that build agents’. But it passed fewer than a quarter of the tests.

The New Stack (AI)
Read full post
Sierra's Hyper-𝜏-bench benchmark evaluates how well AI developer agents can autonomously build other agents. Testing six model and coding harness combinations, Claude Opus 5 achieved the highest success but still passed fewer than 25% of tasks.

More in Agents

Meta Announces Muse AI Agent for Personal Tasks and Organization

Covered by 12 sources
Agents4 min read

Exclusive: Cfo.ai launches an agentic CFO for business founders

SiliconANGLE

A16z Doubles Down On AI Coding Agents Months After Cursor Exit

Forbes