LLM & Text Generation2 min reading time
DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
Apple Research Blog
Read full postResearchers introduced DeepAmbigQA, a dataset of 3,600 multi-hop questions with half involving name ambiguity, to benchmark large language models' ability to provide complete answers. Tests show GPT-5 struggles with answer completeness, scoring low on exact matches for ambiguous and non-ambiguous questions. This highlights the challenge of developing QA systems that effectively resolve ambiguity and integrate multi-step reasoning.




