BrowseComp: A Benchmark for Browsing Agents
Key point
A new benchmark called **BrowseComp** has been released to measure AI agents' ability to deeply explore information.
Details
The importance of AI agents that browse the internet to gather knowledge is growing. Existing benchmarks such as SimpleQA only check basic factual relationships, making them limited in measuring the performance of models that use advanced browsing tools.
Accordingly, a new benchmark called BrowseComp (Browsing Competition) has been released to evaluate the ability to find complex and entangled information on the internet. Consisting of a total of 1,266 challenging problems, this benchmark can be found in OpenAI's simple evals GitHub repository.
BrowseComp utilizes the principle of 'asymmetry of verification.' In other words, it consists of problems that are very difficult to find the answer to, but very easy to verify once found. To ensure the difficulty of the problems, the following criteria were applied:
- Existing models such as GPT-4o and o1 must fail to solve them
- The answer must not be exposed on the first page of search engine results
- They must be difficult for a human to solve within 10 minutes
This benchmark comprehensively evaluates a model's reasoning ability to judge the factuality of internet content, its persistent exploration ability, and its creative search strategies for efficient searching.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.