ConciseSignal

Top AI developer agents pass under 25% in new benchmark

A new open test called Hyper-τ-bench measured how well AI-powered developer agents can build customer service bots on their own. The best-performing system, Anthropic’s Claude Opus 5, completed just 23.9% of tasks correctly, slightly ahead of GPT-5.6 Sol at 22%. Human engineers working with AI scored much higher, hitting 82.2%. The benchmark was designed to be challenging, with none of the six AI setups passing a quarter of the tests.

Why it mattersAI bots still need significant human guidance if you want them to reliably build real-world business tools. These low scores show how far there is to go before you can trust agents to create other agents without oversight.

Sources covering this

The New StackClaude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests.8:14 PM

In this story

Claude
Concise Signal DailyEverything that mattered, every weekday at 7am.

More in AI

18 sources · 3d ago

OpenAI rolls out GPT-6 Astra to most paying users

OpenAI has made its new GPT-6 Astra model available to most paying subscribers a day after its official launch. The rollout, initially described as “messy” by OpenAI’s CEO, now includes ChatGPT Plus, Business, Pro, and Enterprise users, while customers on the lower-priced Go tier do not have access. Astra is also available via the API. OpenAI says the rollout required bringing new systems and additional computing resources online.

6 sources · 1d ago

ChatGPT Images 2.5 adds faster editing tools

OpenAI has updated its ChatGPT Images feature to version 2.5, adding several new editing tools and speeding up image generation. You can now remove backgrounds, resize images to preset dimensions, erase items by brushing over them, and use markup and comment functions directly from a new Edit toolbar. According to TechRadar, the features are available in both web and app versions, though OpenAI hasn't formally announced the rollout yet.

12 sources · 17h ago

Anthropic researcher resigns, warns of AI race risks

Jacob Coxon, who worked on AI training at Anthropic and previously OpenAI, resigned this week. He posted that large AI firms are 'gambling with our lives' by pushing towards systems they may not be able to control. Coxon says both companies know the risks but feel locked in a race to develop powerful AI anyway. He called for drastic steps like a pause on new model improvements.

11 sources · 1d ago

OpenAI says AI cracked landmark math problem, faces credit dispute

OpenAI claims its AI has solved a 200-year-old fluid dynamics equation, one of the Clay Millennium math prize problems. The announcement is under fire: NYU mathematician Tristan Buckmaster says OpenAI rushed their work after learning of his progress and tried to sideline other contributors. OpenAI denies accessing rivals’ data, but both sides accuse the other of trying to shape credit. The company spent millions on computing to finish the proof.