AI Briefing
KOSign in

Microsoft and Hugging Face Release ThinkingBox Benchmark Evaluating AI Agents by Backend State

·2026.10.04 07:56

Key point

The new benchmark evaluates 507 stateful business workflows by verifying final database states rather than tool call success, revealing that 67% of agent failures occur despite clean termination.

1 / 5

Details

Microsoft and Hugging Face have released ThinkingBox, a new benchmark and sandbox environment designed to evaluate AI agents based on their final backend state and side effects, rather than just their final response or valid tool calls. The project addresses a critical gap in agent reliability: while an agent may successfully execute a tool call, the resulting database state may still be incorrect.

Benchmark Structure and Methodology

The benchmark consists of 507 stateful business workflows across five domains: Retail (98), Auto insurance (100), Travel (104), Neobank (104), and Consulting (101). Each task is executed in an isolated MCP tool session, starting from an initialized state with no shared database rows. Agents are evaluated using three metrics:

  • pass@1: Success rate in a single attempt.
  • pass@20: Success in at least one of 20 attempts.
  • Observed 20/20: Success in all 20 attempts.

The evaluation uses a side-effect extractor to identify actual changes made to the backend and a deterministic judge to compare these changes against the required final state. Incorrect, missing, or extra side effects result in failure.

Key Findings: Tool Calls Are Not Outcomes

An ablation study involving 12 LLMs and 121,680 valid trials revealed that 67.24% of failures occurred despite the agent terminating cleanly and calling state-changing tools without final errors. The primary causes of these silent failures were:

  • Wrong field values: 77.61%
  • Unintended extra side effects: 43.30%
  • Missing required side effects: 25.36%

Approximately 80% of all failures were attributed to Tool handling issues rather than reasoning errors, highlighting difficulties in recovering from tool errors, precondition failures, and empty lookups.

Model Performance and Reliability

The benchmark highlights a trade-off between coverage (ability to solve a task at least once) and consistency (ability to solve a task reliably).

  • Top Proprietary Models: Claude Opus 5.5 achieved the highest pass@1 score at 67.16%, followed by Claude Opus 5 (66.50%) and GPT-5.4 (65.36%).
  • Top Open-Weight Model: Kimi-K3 led open-weight models with a pass@1 of 57.37%.
  • Consistency vs. Coverage: Kimi-K3 demonstrated the widest coverage (pass@20 of 93.89%) but low consistency (Observed 20/20: 13.41%). In contrast, Claude Opus 5 had lower coverage (pass@20: 79.09%) but higher consistency (Observed 20/20: 47.53%).
  • Domain Variance: Performance varied significantly by domain, with Retail averaging 59.52% pass@1 compared to just 33.83% for Auto insurance.

Cost Analysis

The study analyzed the cost per successful task attempt and the cost per dependable task (20/20 success rate) using OpenRouter list rates.

  • Cheapest per Single Success: GPT-5.6 Sol at $0.127.
  • Cheapest per Dependable Success: GPT-5.4 at $6.80.
  • Claude Opus 5.5: Costed $0.276 per single success and $7.80 per dependable task.
  • Kimi-K3: Despite lower per-attempt costs, its low consistency resulted in a high cost per dependable task of $20.68.

The data indicates that the model with the lowest cost per single success (GPT-5.6 Sol) is not the most cost-effective for reliable, repeatable operations.

Deployment and Resources

ThinkingBox is available on Hugging Face under the OpenEnv interface. The code is licensed under MIT, the benchmark data under CDLA-Permissive-2.0, and the OpenEnv environment under BSD-3-Clause. The project was developed by the Microsoft Copilot Studio team in collaboration with Toloka, the University of Pittsburgh, Northwestern University, Columbia University, and UC Irvine.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.