🔍 Read the full analysis: Why An AI Agent’s “Done” Doesn’t Always Match The Database on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Microsoft and Hugging Face have made the ThinkingBox benchmark available through Hugging Face. It tests whether AI agents leave business systems in the required state across 507 workflows, with 20 runs per task; its authors report that many attempts failed backend checks even when no final tool error occurred.
Microsoft and Hugging Face have made ThinkingBox, a benchmark for AI agents, available through Hugging Face. Rather than judging only tool calls or final replies, it checks whether an agent leaves a business system in the required state across 507 workflows, each repeated 20 times; the authors report substantial gaps between apparently clean execution and passing those checks, as explored in the original analysis.
ThinkingBox runs agents in isolated sessions using Model Context Protocol (MCP) tools, then checks the backend’s final records and side effects against executable requirements. The workflows span retail, auto insurance, travel, neobanking and consulting. The release says the benchmark is based on the authors’ paper and can also be run through OpenEnv.
In a common-set analysis covering 121,680 valid trials across 12 models, the authors report that 79,853 attempts failed executable checks. Of those failures, 67.24% ended without a final tool error despite the agent invoking a state-changing tool. Among failed attempts, checks identified wrong field values in 77.61%, unintended extra effects in 43.30%, and missing required effects in 25.36%. The categories overlap, so a single failure could involve more than one issue.
The release reports an overall pass@1 score of 67.16% for Claude Opus 5.5 and 57.37% for Kimi-K3, described as the strongest open-weight model in the table. It says Kimi-K3 scored within one point of GPT-6 Astra. These are results from the benchmark authors’ tested setup, not independent measurements of performance across live business systems. The supplied material does not include uncertainty estimates for the model comparisons.
Why Backend State Changes the Score
For companies using agents to handle refunds, support cases, claims or bookings, a fluent response does not establish that the requested work was completed correctly. An agent might close a case prematurely, save an incorrect value, or create an unintended change while appearing to use its tools normally. Checking the recorded outcome tests a different and more direct part of the workflow than evaluating the conversation alone.
The repeated runs also address consistency. Pass@1 measures the share of individual attempts that succeed; pass@20 asks whether a task succeeded at least once in 20 runs; observed 20/20 counts tasks that passed every recorded run. A task that passes all 20 trials provides stronger evidence within this bounded test than one successful attempt, but it does not establish long-term reliability or guarantee the same result in a company’s systems.
As an affiliate, we earn on qualifying purchases.
From Valid Calls to Correct Records
Many agent evaluations focus on whether a system selects tools appropriately or provides a plausible final answer. ThinkingBox’s premise is that those signals can miss a gap between what an agent says it did and what the backend actually records. Its executable checks compare each run’s terminal state and side effects with the task’s requirements.
The release illustrates the distinction with a retail support workflow. An agent investigates a delayed $745 appliance order, opens a ticket and records a timeline. The customer does not qualify for late-delivery compensation under the policy the agent checked, but the carrier exception remains open and the ticket is supposed to stay on hold. The agent instead marks it solved and replies without answering the customer’s underlying question. The authors say the state check fails because the ticket is solved rather than on hold.
“A tool call is not an outcome.”
— Microsoft and Hugging Face, in the ThinkingBox release
business process automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Reported Trials
The reported scores and failure patterns are the authors’ findings on ThinkingBox. The supplied material does not provide the full task specifications, model configurations, or uncertainty estimates needed to assess how precise the reported model differences are. It also does not establish how the tested agents would perform in other organizations’ software, policies or operating conditions.
Passing a task in 20 observed runs is a limited measurement, not proof of durable reliability. Live systems can involve changing records, unusual requests and integrations that may not appear in the benchmark workflows. The source does not say whether the results have been independently replicated, provide a publication date, or establish that the benchmark predicts performance after deployment.
As an affiliate, we earn on qualifying purchases.
Testing Workflows Beyond the Benchmark
The release says researchers and developers can run ThinkingBox through OpenEnv, using isolated MCP tool sessions to examine tasks and compare agent outcomes against executable checks. The supplied material does not identify a future release date or another planned milestone.
For organizations considering agent deployment, the next practical step is to test relevant workflows against their own business rules and inspect both successful and failed runs. ThinkingBox offers a way to examine whether a recorded outcome meets a task’s requirements, but whether its benchmark results translate to a particular organization remains an open question.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does ThinkingBox measure?
It checks whether an AI agent leaves backend records and side effects in the state required by a workflow, rather than relying only on valid tool calls or a plausible final reply.
How many workflows and runs does it include?
The release describes 507 business workflows, each repeated 20 times. The authors also report a common-set analysis of 121,680 valid trials across 12 models.
What did the authors report about failed attempts?
They report that 79,853 attempts in the common-set analysis failed executable checks. Of those failures, 67.24% had no final tool error despite a state-changing tool call. The reported error categories overlap.
Do the scores show how agents will perform in a company’s systems?
No. The scores describe results in the benchmark’s tested setup. The release does not establish that they predict performance across live systems, different policies or long-term use.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
