AI Briefing
KO

EVA-Bench Data 2.0: 3 Domains, 121 Tools, 213 Scenarios

·2026.06.04 21:25

Key point

ServiceNow has released EVA-Bench 2.0, a dataset for evaluating voice AI agents that includes 3 enterprise domains.

Details

EVA-Bench 2.0, designed to evaluate the domain-specific adaptability of voice AI agents, has been released. With this update, the dataset has expanded from a single domain to 3 domains: airline customer service (CSM), enterprise IT service management (ITSM), and healthcare HR service delivery (HRSD).

Key features are as follows:

  • Expanded scale: It includes a total of 213 scenarios and 121 tools, roughly a 4x increase compared to the previous version.
  • Realistic design: It models API schemas and policies used by actual enterprises, covering everything from single-intent and multi-intent calls to adversarial calls that attempt to bypass security.
  • Verified: The solvability of scenarios was verified against the latest frontier models to ensure the fairness of the benchmark.
  • Open source: The entire dataset is available for anyone to download and use via Hugging Face.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.