WANDR Benchmark: Evaluating Research Agents on Wide and Deep Exploration
Key point
Perplexity has released WANDR, a new open benchmark that evaluates the ability to perform broad data collection and deep evidence verification.
Details
Perplexity has unveiled WANDR (Wide ANd Deep Research), an open benchmark containing 500 practical and challenging data-collection tasks for knowledge work. This is an expanded version of DRACO, an existing deep research evaluation tool, and it measures whether an agent can gather vast amounts of information while accurately providing grounding for each piece of information.
The core capabilities of a research agent are divided into two categories. The first is Wide, the ability to find target entities in large numbers, and the second is Deep, the ability to investigate evidence supporting the requested claims for each entity. WANDR deals with complex tasks that require satisfying both of these requirements simultaneously.
WANDR defines tasks using a flexible hierarchical qualification key structure. For example, through a structure like company(n) → employee(m) → url(k), complex exploration paths—finding a certain number of companies, finding employees at each company, and finding evidence URLs for each employee—can be independently verified.
The current state of the art still has a long way to go. According to the evaluation results, even the strongest system only achieved scores of 0.363 soft F1 and 0.133 hard F1, showing that automating wide and deep research remains an unsolved challenge.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.