AI Briefing
KOSign in

ThinkingBox Releases Agent Reliability Benchmark

·2026.10.07 08:00

Key point

A new benchmark has been released that evaluates tasks by running them 20 times and grading based on the final database state.

1 / 3

Details

A research team from Microsoft, the University of Pittsburgh, and Northwestern University has released the ThinkingBox benchmark. Unlike existing benchmarks that only check tool call formats or conversation content, ThinkingBox grades based on the final DB state and side effects after the conversation ends.

Key Finding: One Success Isn't Reliability

The most important finding is the significant gap between succeeding once (pass@k) and succeeding every time (pass^k). They named this the discovery-reliability gap.

  • Kimi-K3: Achieved the highest pass@20 of 93.89% (succeeding at least once in 20 runs), but only reached 17.60% for pass^20 (succeeding in all 20 runs).
  • Claude Opus 5: Had a lower pass@20 of 79.09% compared to Kimi-K3, but achieved the highest overall pass^20 of 47.53%, demonstrating high reliability.
  • GPT-6 Astra: Had a moderate pass@1 of 58.31%, but ranked second with a pass^20 of 46.89%, proving consistent performance.

Evaluation Methodology and Structure

ThinkingBox-Bench includes 507 tasks across 5 business domains: retail, travel/hospitality, auto insurance, neobank IT support, and consulting IT/HR support.

  • Grading Criteria: Does not fix the sequence of correct actions; 1 point is awarded only if all checks pass. There are no partial scores.
  • Repeated Execution: All tasks are executed independently 20 times to calculate pass@1 (average success rate), pass@k, and pass^k metrics.
  • Failure Analysis: 67.24% of failed attempts occurred when the conversation ended normally and there were no tool call errors, yet the final state was inconsistent. Many cases could be misjudged as successful based solely on the response.

Model Performance and Cost Analysis

After evaluating 18 models, it was found that open-weight models do not show a proportional relationship between parameter count and performance. Qwen3.8-27B showed higher performance than DeepSeek-V4-Pro (1.6T).

  • Cost per single success: GPT-5.6-sol was the lowest at $0.127.
  • Cost per task succeeded in all 20 runs: GPT-5.4 was the lowest at $6.80.
  • Reinforcement Learning Fine-tuning (GRPO): Applying GRPO to Qwen3.8-27B improved pass@1 from 51.70% to 60.78%.

Limitations and Operational Recommendations

The communication quality of the final response is not graded, and simulated users follow a cooperative, fixed pattern with no goal changes. The research team recommends verifying the final state, retrying only recoverable errors, and requiring human approval for irreversible changes in actual operations.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.