DeepAmbigQA introduces an automated pipeline to evaluate how well LLMs handle ambiguous multi-hop questions. The benchmark specifically tests a model's ability to distinguish between entities sharing names and integrate evidence across large datasets.
HOW THIS AFFECTS YOU
●
researcherYou can use this to test if your RAG or agentic workflows actually capture all necessary components of a complex query.