AI Briefing
KOSign in

Exa Previews ATLAS Benchmark to Evaluate Agentic Search Accuracy and Cost

·2026.10.08 09:00

Key point

ATLAS uses 547 tasks with a $0.24 grading cost to reveal that no agent achieves a row F1 score above 0.5 under a $1 budget.

1 / 8

Details

Exa has introduced ATLAS (Agentic Tasks for Large Aggregation + Search), a new benchmark designed to evaluate the accuracy and completeness of AI agents on search-intensive workflows. Unlike previous benchmarks that suffer from high memorization rates and saturation, ATLAS focuses on tasks requiring deep web exploration, where frontier models like GPT-6 Astra, GPT-5.6 Sol, and Claude Opus 5 struggle to answer using internal knowledge alone.

Key Findings on Cost and Performance

The benchmark highlights a significant gap in efficient agentic search. No agent achieved a row F1 score above 0.5 when executed with a budget under $1. Even the most expensive search agents missed approximately one-third of the golden results, indicating substantial room for improvement in completeness. When the model harness is fixed, the choice of search backend significantly impacts performance, with Exa defining the cost-performance Pareto frontier and showing up to a 16% score difference depending on the backend used.

Benchmark Design and Metrics

ATLAS consists of 547 deep and wide research tasks generated from anonymized search demand clusters. Each task requires discovering all entities meeting specific conditions and enriching them with 2-10 multi-hop attributes. The evaluation uses three F1 metrics:

  • Discovery F1: Measures entity discovery only.
  • Item F1: Measures cell-level accuracy (discovery + enrichment).
  • Row F1: The headline metric, requiring every cell in a row to be correct.

Grading is highly efficient and deterministic, costing approximately $0.24 to grade all 547 tasks for a single system. This is significantly cheaper than alternatives like WANDR, which costs $0.88 per answer and takes 7.6 minutes per answer, compared to ATLAS's $0.0004 and 15 seconds per answer.

Validation and Future Updates

The dataset was constructed using a multi-stage pipeline involving frontier models and independent verification, achieving a gold value error rate below 0.9%. ATLAS is designed to be continuously updated, with plans to regenerate the dataset every 3-6 months based on fresh search demand. Exa intends to release the tasks, golden tables, and grader details in the coming weeks.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.