AgentsMachine Learning9 min reading time

Claude performed best on a new benchmark for ‘agents that build agents’. But it passed fewer than a quarter of the tests.

The New Stack (AI)
Read full post
Sierra's Hyper-𝜏-bench benchmark evaluates how well AI developer agents can autonomously build other agents. Testing six model and coding harness combinations, Claude Opus 5 achieved the highest success but still passed fewer than 25% of tasks.

More in Agents

Meta Announces Muse AI Agent for Personal Tasks and Organization

Covered by 11 sources
Agents5 min read

Abacus.AI Releases Three Open-Weight Smaug Models for Agentic Workloads

Unite.AI

Winmau And Autodarts Bring Smart Scoring To Your Dumb Dartboard

Forbes