[arXiv]score: 0.24
From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models
August 12, 2026
MPAR-Bench evaluates reasoning breadth by requiring models to integrate semantically diverse clues into a single coherent answer. This bilingual English-Chinese benchmark utilizes 1,000 multi-agent generated items to isolate parallel associative reasoning from traditional linear depth. The dataset employs embedding-based diversity filtering to ensure clue sets provide distinct semantic paths.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy