MiGUE-Bench provides a systematic evaluation of LLM performance across event detection, relation reasoning, structure induction, and future prediction. The framework utilizes the MiGUE-Pipeline, an LLM-driven self-correcting annotation method, to generate scalable, high-quality event data across varying document granularities.
HOW THIS AFFECTS YOU
●
researcherYou can use this benchmark to assess how well models handle complex event extraction across multiple document scales.