AgentsMachine Learning9 min reading time

Claude performed best on a new benchmark for ‘agents that build agents’. But it passed fewer than a quarter of the tests.

The New Stack (AI)
Read full post
Sierra's Hyper-𝜏-bench benchmark evaluates how well AI developer agents can autonomously build other agents. Testing six model and coding harness combinations, Claude Opus 5 achieved the highest success but still passed fewer than 25% of tasks.

More in Agents

Meta Announces Muse AI Agent for Personal Tasks and Organization

Covered by 11 sources
Agents4 min read

Amazon makes its agentic AI platform Quick generally available for desktop on Windows and macOS

Covered by 2 sources
Agents5 min read

Abacus.AI Releases Three Open-Weight Smaug Models for Agentic Workloads

Unite.AI