Microsoft Releases ThinkingBox: 507 Stateful Workflows Reveal Gap Between Agent Discovery and Repeatability
Key point
The ThinkingBox benchmark evaluates 9 models across 10,140 trials per model, showing that high pass@20 scores do not guarantee consistent task completion.
Details
Microsoft has made public ThinkingBox, a benchmark comprising 507 policy-conditioned business workflows across five domains (retail, travel/hospitality, auto insurance, neobank internal IT, and consulting IT/HR). The benchmark assesses agent reliability by running each task 20 times from a clean backend state, resulting in 10,140 trials per model.
Key Findings on Reliability
The evaluation reveals a significant divergence between an agent's ability to solve a task at least once (pass@20) and its ability to solve it consistently (all-20):
- Kimi-K3 achieved the highest coverage, solving 93.89% of tasks at least once, but only 13.41% on all 20 attempts.
- Claude Opus 5 solved 79.09% of tasks at least once but demonstrated higher repeatability, succeeding on all 20 attempts for 47.53% of tasks.
- Qwen3.8-27B solved 89.35% of tasks at least once but only 7.50% consistently.
Failure Analysis
A retrospective ablation over 121,680 valid trials showed that 67.24% of failed attempts terminated cleanly without tool errors, meaning traditional completion proxies would incorrectly mark them as successful. Among these clean-terminating failures, the state checks identified the following primary failure modes (categories overlap):
- Wrong field values: 77.61%
- Unintended extra effects: 43.30%
- Missing required effects: 25.36%
Availability
The benchmark, code, and dataset are publicly available on GitHub and Hugging Face. The environment is accessible via HF OpenEnv, allowing developers to evaluate their own models against the 507 tasks.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.