DeepAmbigQA: Ambiguous Multi-Hop Questions for Evaluating LLM Answer Completeness
Key point
DeepAmbigQA combines ambiguity and multi-hop reasoning to evaluate the answer completeness of LLMs.
Details
DeepAmbigQA is a dataset that evaluates how completely LLMs using search tools gather answers to complex questions. It addresses questions that require distinguishing between multiple entities with the same name and collecting and integrating information from large candidate sets.
The researchers built the DEEPAMBIGQAGEN automated generation pipeline based on text corpora and linked knowledge graphs. This enabled the systematic generation of natural and verifiable questions that incorporate name ambiguity and multi-hop reasoning.
DEEPAMBIGQA consists of 3,600 questions in total, half of which require explicit name disambiguation. In experiments, even the latest model GPT-5 struggled to complete answers, with exact match scores reaching only 0.13 for ambiguous questions and 0.21 for non-ambiguous questions.
The research results demonstrate the need to evaluate and improve the ability to exhaustively collect and integrate multiple pieces of information, beyond simple retrieval or single-answer prediction.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.