Perplexity Releases WANDR Benchmark
Key point
Perplexity has released WANDR, an open benchmark that evaluates a research agent's ability to collect data at scale.
Details
WANDR (Wide ANd Deep Research) is an open benchmark released by Perplexity, consisting of 500 data-collection tasks drawn from real knowledge work performed by research agents. It covers realistic tasks such as competitor analysis, due diligence, market analysis, and talent sourcing, and each task requires discovering anywhere from dozens to thousands of independently verifiable records.
The core question of the benchmark is whether an agent can perform both Wide and Deep research at the same time.
- Wide: the ability to exhaustively discover a large set of entities that satisfy given conditions
- Deep: the ability to back every discovered entity with factual evidence
Tasks are represented as a Qualification Key Hierarchy tree structure. For example: company(n) → employee(m) → url(k). This structure can express everything from flat lists to matrix-style tasks.
According to the evaluation results, even the strongest system only achieved Soft F1 0.363 / Hard F1 0.133, leading Perplexity to conclude that "wide and deep research is still far from being solved." The benchmark comprises a total of 170,495 records, with an average of 80 members per task.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.