AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why An AI Agent’s “Done” Doesn’t Always Match The Database on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Microsoft and Hugging Face have made the ThinkingBox benchmark available through Hugging Face. It tests whether AI agents leave business systems in the required state across 507 workflows, with 20 runs per task; its authors report that many attempts failed backend checks even when no final tool error occurred.

Microsoft and Hugging Face have made ThinkingBox, a benchmark for AI agents, available through Hugging Face. Rather than judging only tool calls or final replies, it checks whether an agent leaves a business system in the required state across 507 workflows, each repeated 20 times; the authors report substantial gaps between apparently clean execution and passing those checks, as explored in the original analysis.

ThinkingBox runs agents in isolated sessions using Model Context Protocol (MCP) tools, then checks the backend’s final records and side effects against executable requirements. The workflows span retail, auto insurance, travel, neobanking and consulting. The release says the benchmark is based on the authors’ paper and can also be run through OpenEnv.

In a common-set analysis covering 121,680 valid trials across 12 models, the authors report that 79,853 attempts failed executable checks. Of those failures, 67.24% ended without a final tool error despite the agent invoking a state-changing tool. Among failed attempts, checks identified wrong field values in 77.61%, unintended extra effects in 43.30%, and missing required effects in 25.36%. The categories overlap, so a single failure could involve more than one issue.

The release reports an overall pass@1 score of 67.16% for Claude Opus 5.5 and 57.37% for Kimi-K3, described as the strongest open-weight model in the table. It says Kimi-K3 scored within one point of GPT-6 Astra. These are results from the benchmark authors’ tested setup, not independent measurements of performance across live business systems. The supplied material does not include uncertainty estimates for the model comparisons.

At a glance
reportWhen: Availability announced in source materi…
The developmentMicrosoft and Hugging Face have released ThinkingBox, a benchmark that evaluates AI agents by checking database states and side effects after business workflows.
At a glance
announcementWhen: Now available through Hugging Face; the…
The developmentMicrosoft and Hugging Face released ThinkingBox through Hugging Face, a benchmark for evaluating AI agents by their backend changes across repeated workflow trials.

Why Backend State Changes the Score

For companies using agents to handle refunds, support cases, claims or bookings, a fluent response does not establish that the requested work was completed correctly. An agent might close a case prematurely, save an incorrect value, or create an unintended change while appearing to use its tools normally. Checking the recorded outcome tests a different and more direct part of the workflow than evaluating the conversation alone.

The repeated runs also address consistency. Pass@1 measures the share of individual attempts that succeed; pass@20 asks whether a task succeeded at least once in 20 runs; observed 20/20 counts tasks that passed every recorded run. A task that passes all 20 trials provides stronger evidence within this bounded test than one successful attempt, but it does not establish long-term reliability or guarantee the same result in a company’s systems.

Amazon

AI workflow testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Valid Calls to Correct Records

Many agent evaluations focus on whether a system selects tools appropriately or provides a plausible final answer. ThinkingBox’s premise is that those signals can miss a gap between what an agent says it did and what the backend actually records. Its executable checks compare each run’s terminal state and side effects with the task’s requirements.

The release illustrates the distinction with a retail support workflow. An agent investigates a delayed $745 appliance order, opens a ticket and records a timeline. The customer does not qualify for late-delivery compensation under the policy the agent checked, but the carrier exception remains open and the ticket is supposed to stay on hold. The agent instead marks it solved and replies without answering the customer’s underlying question. The authors say the state check fails because the ticket is solved rather than on hold.

“A tool call is not an outcome.”

— Microsoft and Hugging Face, in the ThinkingBox release

Amazon

business process automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Reported Trials

The reported scores and failure patterns are the authors’ findings on ThinkingBox. The supplied material does not provide the full task specifications, model configurations, or uncertainty estimates needed to assess how precise the reported model differences are. It also does not establish how the tested agents would perform in other organizations’ software, policies or operating conditions.

Passing a task in 20 observed runs is a limited measurement, not proof of durable reliability. Live systems can involve changing records, unusual requests and integrations that may not appear in the benchmark workflows. The source does not say whether the results have been independently replicated, provide a publication date, or establish that the benchmark predicts performance after deployment.

Amazon

database state verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Testing Workflows Beyond the Benchmark

The release says researchers and developers can run ThinkingBox through OpenEnv, using isolated MCP tool sessions to examine tasks and compare agent outcomes against executable checks. The supplied material does not identify a future release date or another planned milestone.

For organizations considering agent deployment, the next practical step is to test relevant workflows against their own business rules and inspect both successful and failed runs. ThinkingBox offers a way to examine whether a recorded outcome meets a task’s requirements, but whether its benchmark results translate to a particular organization remains an open question.

Amazon

AI model performance benchmarking

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does ThinkingBox measure?

It checks whether an AI agent leaves backend records and side effects in the state required by a workflow, rather than relying only on valid tool calls or a plausible final reply.

How many workflows and runs does it include?

The release describes 507 business workflows, each repeated 20 times. The authors also report a common-set analysis of 121,680 valid trials across 12 models.

What did the authors report about failed attempts?

They report that 79,853 attempts in the common-set analysis failed executable checks. Of those failures, 67.24% had no final tool error despite a state-changing tool call. The reported error categories overlap.

Do the scores show how agents will perform in a company’s systems?

No. The scores describe results in the benchmark’s tested setup. The release does not establish that they predict performance across live systems, different policies or long-term use.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Deepfakes 101: How AI-Generated Videos Are Changing What We See

Keen to understand how AI-generated deepfakes are transforming visual media and what risks they pose? Keep reading to uncover the full picture.

Why 3D Printing Accessories Matter More Than New Users Expect

What you choose to add to your 3D printer can dramatically improve results, making all the difference in your printing journey—find out why.

XAI Grok 4.6: Pushing AI Boundaries While Cutting Costs Significantly

xAI announces Grok 4.6, claiming near-frontier performance with 85% lower costs, but details on benchmarks, release, and comparison remain unclear.

The Tech Behind ByteDance’s 10 Trillion Parameter AI Model And Its GPU Power

ByteDance reportedly plans to train a 10 trillion-parameter AI model using around 30,000 GPUs, but has not officially confirmed the project. Key details remain unknown.