Microsoft and Hugging Face launch ThinkingBox agent test
Microsoft and Hugging Face have released ThinkingBox, a tool that evaluates AI agents by checking not just their actions but whether they achieve the correct outcome in backend databases. In tests spanning over 121,000 workflow attempts with 12 AI models, nearly two-thirds failed to land in the intended database state, despite carrying out valid steps. ThinkingBox is now openly available for research use.
- ThinkingBox grades agents on end-state in databases
- Most agent runs failed executable outcome checks
- Tested 121,000 cases across 12 LLM models
- Tool now available through Hugging Face
Sources covering this
In this story
More in AI
OpenAI safety leader resigns, warns of 'broken' AI culture
David Robinson, who helped shape safety reports for OpenAI’s biggest AI releases, has left the company, publicly warning that its…
Google testing expanded Gemini access on Macs
Google is quietly testing a feature that could let its Gemini AI interact with any file, app, and some web actions on Macs—no need for…
Anthropic adds always-on agent abilities to Claude
Anthropic has quietly added always-on AI agent features to its Claude assistant, letting it handle ongoing or scheduled tasks even after…
Custom software links iPhone to MacBook for faster AI model runs
A developer used custom open-source software, 'backburner,' to split an AI workload between an iPhone 17 Pro Max and a MacBook Pro with…