AI Briefing
KO
Pick

ITBench-AA: Enterprise IT Agent Benchmark Released

·2026.05.28 02:20

Key point

Artificial Analysis and IBM have launched ITBench-AA, a new benchmark for evaluating enterprise IT agent performance.

1 / 2

Details

Artificial Analysis and the IBM Software Innovation Lab have announced a new benchmark, ITBench-AA, which evaluates agentic enterprise IT tasks, particularly SRE (Site Reliability Engineering) capabilities.

The benchmark covers Kubernetes incident response, including high-difficulty tasks that require models to analyze logs and trace dependencies to identify root causes within complex infrastructure.

Key findings:

  • Frontier model performance: All frontier models scored below 50%, suggesting that agentic IT tasks remain a highly challenging domain.
  • Model rankings: Claude Opus 4.7 showed the highest performance at 47%, followed by GPT-5.5 (46%) and Qwen3.7 Max (42%).
  • Open-weight models: GLM-5.1 (Reasoning) led with 40%, followed by DeepSeek V4 Pro (38%) and Gemma 4 31B (37%).
  • Efficiency issues: Simply taking more steps (turns) did not lead to improved accuracy. In fact, excessive investigation was found to tend to cause false positives.

The benchmark's scope is planned to expand to FinOps and CISO-related tasks in the future.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.